Visual model upgrading method, electronic equipment, storage medium and program product
By automatically generating and verifying training samples, the problems of low accuracy and resource consumption of traditional AI vision models in different scenarios are solved, and efficient automation upgrade and accuracy improvement of the model is achieved.
Patent Information
- Application Number
- CN202510865175.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Traditional AI vision models are difficult to maintain high accuracy and efficiency in different scenarios, and each training and iteration upgrade requires manual collection and labeling of pictures, which consumes a lot of manpower and material resources. There are few training samples, resulting in low model accuracy.
By obtaining the image samples and detection results uploaded by the target terminal, generating training samples and updating the database, using the pre-trained language model to verify the sample accuracy, automatically identify the model update time, and perform model training and upgrade when the training sample reaches the threshold, and replacing the manual verification process with the trained model.
It improves the accuracy of model recognition, saves labor costs, improves training efficiency, and realizes the automated iteration and upgrading of the model and continuous learning.
Smart Images

Figure CN120371356A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method for upgrading a vision model, an electronic device, a storage medium, and a program product. Background Art
[0002] With the development of artificial intelligence technology, more and more places have started to carry out intelligent transformation and deploy intelligent monitoring systems. These systems cover application scenarios such as smoke and fire detection, personnel aggregation detection, and electronic fences, aiming to improve safety and operation efficiency. The AI (Artificial Intelligence) model, as the basis of deep learning applications, such as CNN (Convolutional Neural Network) and RNN (Recurrent Neural Network), is of great importance.
[0003] However, due to differences in environment, data characteristics, and device performance among different places, traditional single models are difficult to maintain high accuracy and efficiency in all scenarios. Therefore, continuous training and iterative upgrading of the model have become essential functions of the edge management software platform. In the training process of traditional AI vision models, this process requires manual collection and annotation of pictures in different scenarios, and then training and upgrading are completed. Each iterative upgrade consumes a large amount of manpower and material resources, and this iteration is continuous, and it is often very difficult to obtain sufficient training samples. In the actual production environment, the initially deployed model usually cannot meet the application requirements due to insufficient samples, resulting in the subsequent need for model training and upgrading. Summary of the Invention
[0004] The present invention provides a method for upgrading a vision model, an electronic device, a storage medium, and a program product, so as to at least solve the technical problems in the related art that each training and iterative upgrade requires manual collection and annotation of pictures, consumes a large amount of manpower and material resources, and the small number of training samples leads to low model accuracy.
[0005] The present invention provides a method for upgrading a vision model. The vision model is deployed on a target terminal, and the vision model outputs the picture detection result of a picture sample. The method is applied to a server and includes: obtaining the picture sample and the picture detection result uploaded by the target terminal; generating a training sample according to the picture sample and the picture detection result, updating a database according to the training sample, and identifying the update time of the vision model; obtaining the number of newly added training samples in the database since the update time, and when the number of training samples is greater than a number threshold, training the vision model using the training samples in the database, and upgrading the vision model of the target terminal using the trained vision model.
[0006] The present invention also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above visual model upgrade methods when executing the computer program.
[0007] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any of the above visual model upgrade methods when executed by a processor.
[0008] The present invention also provides a computer program product including a computer program, which implements the steps of any of the above visual model upgrade methods when executed by a processor.
[0009] Through the present invention, since training samples can be generated according to picture samples and picture detection results, the database can be updated using the training samples, the update time of the visual model can be identified, and when the number of training samples is greater than the quantity threshold, the visual model in the database is trained using the training samples, and the visual model of the target terminal is upgraded using the trained visual model, so that the number of training samples reaches the desired quantity threshold, and the visual model can be used for picture parsing to replace the manual verification process and automatically perform model training. Therefore, the technical problems in the related art such as low model accuracy due to few training samples, and the need for manual picture collection and annotation for each training and iterative upgrade, consuming a large amount of manpower and material resources can be solved, and technical effects such as improving model recognition accuracy, saving labor costs, and enhancing training efficiency can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] To more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0011] Figure 1 It is a schematic flowchart of a visual model upgrade method provided by an embodiment of the present invention; Figure 2 It is a system flowchart provided by an embodiment of the present invention; Figure 3 It is a schematic flowchart of a multi-modal pre-trained language model sample verification process provided by an embodiment of the present invention; Figure 4 It is a schematic flowchart of a multi-modal pre-trained language model negative sample enhancement process provided by an embodiment of the present invention; Figure 5 It is a schematic block diagram of a visual model upgrade device provided by an embodiment of the present invention; Figure 6A schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0012] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0013] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0014] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0015] Figure 1 A schematic flowchart of a method for upgrading a visual model provided by an embodiment of the present invention. The visual model is deployed on a target terminal, and the visual model outputs the image detection result of an image sample. The method is applied to a server, as Figure 1 shown, the method specifically includes the following steps: In step S101, an image sample and an image detection result uploaded by the target terminal are obtained.
[0016] Among them, the image sample is an image collected by the target terminal and containing an object to be recognized; the image detection result is the result obtained after analyzing the image sample, including but not limited to object labels and coordinate box information.
[0017] It can be understood that the embodiments of the present invention collect image samples with preliminary recognition information and their corresponding image detection results, such as object labels and position information, from target terminals deployed in various application scenarios as the basis for further processing and analysis in the following steps.
[0018] In step S102, a training sample is generated according to the image sample and the image detection result, the database is updated according to the training sample, and the update time of the visual model is identified.
[0019] Among them, the visual model refers to an artificial intelligence model specifically used for processing and analyzing image data, usually constructed based on deep learning technologies such as CNN, which can extract features from images and perform various tasks, such as object recognition, classification, detection, segmentation, etc.; the update time of the recognized visual model refers to the time point when the visual model completes an iterative upgrade and is ready.
[0020] It can be understood that the embodiments of the present invention use the picture samples and picture detection results collected and uploaded by the target terminal to generate qualified training samples, and use these training samples to update the database. During this process, the exact time of this iterative upgrade of the visual model is recorded for subsequent tracking and management of the model version.
[0021] In the embodiments of the present invention, updating the database according to the training samples further includes: verifying the training samples and updating the database using the verification results.
[0022] It can be understood that before using the training samples to update the database, the embodiments of the present invention need to verify these samples first and update the database using the verification results to ensure their accuracy and reliability. Among them, verification is the process of checking the accuracy of the generated training samples, which is specifically as follows: In the embodiments of the present invention, verifying the training samples includes: obtaining the picture samples and picture detection results in the training samples; extracting the coordinate box information in the picture detection results, and using the coordinate box information to crop the picture samples to obtain new pictures; generating model verification questions according to the type of the visual model, and calling the pre-trained language model on the server according to the model verification questions and the new pictures; verifying the training samples according to the feedback results of the pre-trained language model.
[0023] Among them, the model verification questions are used to inquire about the objects and their positions in the pictures, etc., for verifying the picture samples; the pre-trained language model is a high-performance, large-scale multi-modal machine learning model deployed on the server, which can understand the content of the pictures and answer questions about the pictures.
[0024] It can be understood that when verifying the training samples in the embodiments of the present invention, first obtain the picture samples and their detection results in the sample, extract the coordinate box information from them, use this coordinate box information to crop the original picture samples to generate new pictures, construct corresponding model verification questions according to the specific type of the visual model, and use these questions together with the newly generated pictures as inputs to call the pre-trained language model on the server. Evaluate and correct the errors in the training samples according to the feedback results returned by the pre-trained language model to ensure the accuracy and reliability of the training samples, thereby improving the quality of the final visual model, greatly improving the data processing efficiency in an automated manner, and reducing the need for manual intervention.
[0025] In an embodiment of the present invention, cropping a picture sample using coordinate frame information to obtain a new picture includes: determining a picture frame of the picture sample according to the coordinate frame information; cropping the image in the area of the picture frame, and placing the area image on a picture template to obtain a new picture, where the position of the area image on the picture sample and the picture template is the same.
[0026] Wherein, the picture template is a new background picture having the same size as the original picture, and is used to place the cropped area image at a specific position, so as to generate a new composite picture.
[0027] It can be understood that in the embodiment of the present invention, cropping a picture sample using coordinate frame information to obtain a new picture, first determines the position of a specific object in the picture sample according to the coordinate frame information, that is, determines the picture frame of the picture sample, and then, crops the corresponding area image according to the coordinate frame, and places the area image at the corresponding position on a background picture template having the same size as the original picture, that is, the original coordinate frame position, so that the position of the cropped object in the new picture is the same as its position in the original picture, creating a new picture focused on a specific object in the original picture while keeping its spatial relationship in the original picture unchanged.
[0028] In an embodiment of the present invention, verifying a training sample according to the feedback result of a target pre-trained language model includes: identifying whether the feedback result is empty; if the feedback result is empty, deleting the coordinate frame information of the corresponding picture sample; if the feedback result is not empty, comparing whether the label information in the picture detection result and the feedback result is consistent, if the label information is inconsistent, deleting the coordinate frame information of the corresponding picture sample, if the label information is consistent, calculating new coordinate frame information of the picture sample according to the coordinate frame information in the picture detection result and the feedback result.
[0029] It can be understood that the method for verifying a training sample according to the feedback result of a target pre-trained language model in the embodiment of the present invention is: first checking whether the feedback result is empty, if the feedback result is empty, it indicates that the pre-trained language model fails to identify any object, and at this time, the coordinate frame information and label in the corresponding picture sample should be deleted; if the feedback result is not empty, further comparing whether the label information in the original detection result of the picture sample and the feedback result of the pre-trained language model is consistent, if the two label information is inconsistent, it means there is an identification error, and at this time, the coordinate frame information in this picture sample also needs to be deleted; on the contrary, if the label information is consistent, combining the picture detection result and the coordinate frame information provided by the pre-trained language model, using the above weighted formula to calculate the new coordinate frame position, so as to update the coordinate frame information in the picture sample, thus ensuring the accuracy and reliability of the training sample and providing high-quality data support for subsequent model training.
[0030] In an embodiment of the present invention, calculating new coordinate box information of a picture sample according to the coordinate box information in the picture detection result and the feedback result includes: obtaining the respective weights of the visual model and the pre-trained language model; calculating the new coordinate box information of the picture sample by weighted calculation according to the coordinate box information in the picture detection result, the coordinate box information in the feedback result, and the respective weights of the visual model and the pre-trained language model.
[0031] Wherein, the weight coefficient is a constant, which is specifically set according to actual needs and is not specifically limited herein; calculating the new coordinate box information of the picture sample according to the coordinate box information in the picture detection result and the feedback result is implemented through the formula as follows. To calculate the new coordinates of the detection box, and are weight coefficients, and both are constants. For example, take . is the coordinate returned by the pre-trained language model, is the detection coordinate uploaded by the visual model.
[0032] It can be understood that in order to calculate the new coordinate box information of the picture sample in the embodiment of the present invention, it is first necessary to obtain the respective weights of the visual model and the pre-trained language model, and use the coordinate box information in the picture detection result, the coordinate box information in the feedback result of the pre-trained language model, and the weights of these two models to obtain a more accurate new coordinate box through weighted calculation. Specifically, this process fuses the coordinate box detected by the visual model and the coordinate box provided by the pre-trained language model according to a pre-set weight ratio, so as to obtain an optimized new coordinate box information, thereby integrating the advantages of the two models, reducing the errors that may occur in a single model, and finally improving the accuracy of the new coordinate box information.
[0033] In an embodiment of the present invention, updating the database using the verification result includes: using the picture sample with the coordinate box information deleted as a negative sample; using the picture sample carrying the new coordinate box information and the label information as a positive sample; adding the negative sample and the positive sample to the database.
[0034] It can be understood that when updating the database using the verification result in the embodiment of the present invention, first, the picture samples whose coordinate box information has been deleted due to recognition errors are marked as negative samples. At the same time, the picture samples that have been verified and carry the updated coordinate box information and label information are used as positive samples, and these two types of samples are added to the database together, ensuring that the training samples in the database include both positive samples that can help the model learn correct features and negative samples that can help the model avoid false alarms, thereby supporting more efficient and accurate model training and iterative upgrading, improving the performance of the finally deployed visual model, and enabling it to maintain high accuracy and reliability in various complex environments.
[0035] In step S103, obtain the number of training samples newly added to the database since the update time. When the number of training samples is greater than the quantity threshold, use the training samples in the database to train the visual model, and use the trained visual model to upgrade the visual model of the target terminal.
[0036] Among them, the quantity threshold is a preset value. When the number of newly added training samples exceeds this threshold, the training process of the visual model is triggered, which is specifically set according to the actual situation and will not be specifically limited here; the trained visual model upgrades the visual model of the target terminal by deploying this newly trained and optimized model to the target terminal to replace the original old version visual model.
[0037] It can be understood that the embodiment of the present invention obtains the number of training samples newly added since the last update time of the database. Once it is found that the number of newly added training samples exceeds the preset quantity threshold, the process of using these newly added training samples to train the visual model is started. After the training is completed, a visual model with better performance is obtained. Deploy this newly trained and optimized model to each target terminal to replace the original old version visual model, realizing the automatic iterative upgrade of the model, ensuring that the visual model on the edge device can continuously learn the latest data, thereby continuously improving the recognition accuracy and efficiency, adapting to the changing actual application scenario requirements, and greatly improving the efficiency of model maintenance and update in an automated manner.
[0038] In the embodiment of the present invention, the training samples include positive samples and negative samples. Using the training samples in the database to train the visual model includes: selecting positive samples and negative samples from the database; generating a training data set according to the positive samples and negative samples; using the training data set to train the visual model.
[0039] It can be understood that in the embodiment of the present invention, the visual model is trained using the training samples in the database. First, positive samples and negative samples are selected from the database. The positive samples provide the correctly labeled target object and location information to help the model learn to recognize the correct object; while the negative samples show non-target objects or scenes that are prone to misjudgment to enhance the model's ability to distinguish the background from the target. A training data set is generated according to the selected positive samples and negative samples. This data set covers various situations that the model needs to learn. Using this comprehensive training data set to train the visual model enables the model to continuously adjust its own parameters, so as to more accurately perform the target recognition task in actual applications.
[0040] In the embodiment of the present invention, before using the training data set to train the visual model, it further includes: enhancing the number of negative samples in the training data set based on the training samples in the database.
[0041] It can be understood that before training the vision model using the training dataset in the embodiments of the present invention, it is necessary to first increase the number of negative samples in the training dataset based on the training samples in the database, so that the training dataset is more comprehensive, helping the trained vision model to more effectively distinguish the background from non-target objects, reducing the false alarm rate, and thus improving the overall performance and accuracy of the model. The specific process is as follows: In the embodiments of the present invention, increasing the number of negative samples in the training dataset based on the training samples in the database includes: obtaining the number of negative samples in the training dataset; if the number of negative samples is less than the target number, calculating the difference number between the number of negative samples and the target number; randomly selecting training samples from the database based on the difference number, and increasing the number of negative samples in the training dataset based on the training samples.
[0042] Among them, the target number is specifically set according to actual needs and will not be specifically limited here; increasing the number of negative samples in the training dataset based on the training samples will be described in detail below and will not be elaborated here.
[0043] It can be understood that when preparing the training data in the embodiments of the present invention, it is first necessary to check whether the number of negative samples therein has reached the preset target number. If it is found that the existing number of negative samples is insufficient, the required difference number is calculated. Based on this difference number, a certain number of training samples are randomly selected from the database, and the number of negative samples in the training dataset is increased using these samples, ensuring that the training data has sufficient diversity, which helps to improve the learning effect and recognition accuracy of the vision model, enabling it to more accurately distinguish non-target objects and reducing the false alarm rate.
[0044] In the embodiments of the present invention, randomly selecting training samples from the database based on the difference number includes: if the difference number is less than the preset number, randomly selecting the preset number of training samples from the database; if the difference number is greater than or equal to the preset number, randomly selecting the difference number of training samples from the database.
[0045] Among them, the preset number is specifically set according to the actual situation and will not be specifically limited here.
[0046] It can be understood that when the embodiments of the present invention select training samples from the database according to the difference number to increase negative samples, it is first determined whether the difference number is less than the preset number. If the difference number is small, that is, the number of negative samples to be increased is lower than the preset value, the difference number is directly set as the preset number, and the preset number of training samples are randomly selected from the database; on the contrary, if the difference number is greater than or equal to the preset number, it means that more negative samples need to be increased, then the difference number of training samples is randomly selected from the database for enhancement, ensuring that the negative sample quantity can be effectively increased to improve the quality of model training.
[0047] In an embodiment of the present invention, enhancing the number of negative samples in the training dataset based on training samples includes: extracting sample images from the training samples; inputting the sample images into a pre-trained language model of a server, and using the image generation function of the pre-trained language model to generate images without label information; generating negative samples according to the image description information and the images without label information.
[0048] It can be understood that in the embodiment of the present invention, to increase the number of negative samples in the training dataset, first, sample images are extracted from the selected training samples, these sample images are input into the pre-trained language model on the server, and its image generation function is used to generate new images without any label information. Combining specific image description information, the generated new images are used to create negative samples, thereby enriching the content of the training dataset, improving the ability of the visual model to distinguish non-target objects, reducing the false alarm rate, and ultimately enhancing the accuracy and reliability of the model.
[0049] In an embodiment of the present invention, before generating negative samples according to the image description information and the images without label information, it further includes: obtaining misjudgment data of the visual model; generating image description information according to the misjudgment data.
[0050] Among them, the misjudgment data of the visual model refers to the data where the current visual model wrongly identifies a positive sample as a negative sample, or wrongly marks a negative sample as containing a target object, which is obtained by comparing the model prediction result with the actual situation.
[0051] It can be understood that in the embodiment of the present invention, before generating negative samples using the images without label information and the corresponding description information, it is necessary to obtain the misjudgment data generated by the current visual model. This includes the situation where the model wrongly marks an image that actually does not contain a target object as a target object, or fails to correctly identify the existence of a target object. According to the collected misjudgment data, detailed image description information is generated. This description information will specifically point out the key background conditions or environmental factors that are likely to cause model misjudgment. The purpose of doing this is to specifically generate new negative sample images, thereby helping to train a more accurate visual model, reducing the possibility of similar misjudgments in the future, and contributing to improving the overall performance and accuracy of the model.
[0052] In an embodiment of the present invention, after enhancing the number of negative samples in the training dataset based on training samples, it further includes: obtaining the total number of training samples in the training dataset; setting an enhancement stop threshold according to the total number of training samples, where the enhancement stop threshold is less than the total number of training samples; if the number of negative samples is less than the enhancement stop threshold, stop enhancing the number of negative samples in the training dataset, and if the number of negative samples is greater than or equal to the enhancement stop threshold, delete some negative samples in the training dataset.
[0053] Among them, the total number of training samples refers to the number of all training samples contained in the entire training dataset, including positive samples and negative samples; the augmentation stop threshold is a preset value used to determine when to stop adding new negative samples to the training dataset, which is specifically set according to the actual situation and will not be specifically limited here; the number of negative samples refers to the number of existing negative samples in the current training dataset.
[0054] It can be understood that after augmenting the number of negative samples in the training dataset based on the training samples in the embodiments of the present invention, it is necessary to obtain the total number of all training samples in the training dataset, and set an augmentation stop threshold according to this total number. This threshold is less than the total number of training samples and is used to control the increase amount of negative samples. Check the number of negative samples in the current training dataset. If it is found that the number of negative samples is less than the set augmentation stop threshold, then stop augmenting the number of negative samples in the training dataset and end the negative sample augmentation; conversely, if the number of negative samples has reached or exceeded the augmentation stop threshold, then stop adding new negative samples and it is necessary to randomly delete some negative samples from the existing training dataset to ensure that the number of negative samples does not exceed the set threshold, so as to maintain the balance between positive and negative samples in the training dataset and avoid affecting the training effect of the model due to an excessive number of a certain type of sample, thereby ensuring that the finally trained visual model has good generalization ability and recognition accuracy.
[0055] According to the visual model upgrade method proposed by the embodiments of the present invention, training samples can be generated based on picture samples and picture detection results, the database can be updated using the training samples, the update time of the visual model can be identified. When the number of training samples is greater than the quantity threshold, the visual model is trained using the training samples in the database, and the visual model of the target terminal is upgraded using the trained visual model, so that the number of training samples reaches the desired quantity threshold, and the visual model can be used for picture parsing to replace the manual verification process, and model training is automatically performed, achieving technical effects such as improving the model recognition accuracy, saving labor costs, and enhancing the training efficiency.
[0056] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.
[0057] The visual model upgrade method will be further described below through a specific embodiment.
[0058] This embodiment proposes an AI model automatic training and upgrade method under an edge distributed system. This method requires the cooperation of the edge side and the central side (cloud side) to achieve cloud-edge cooperation.
[0059] This method is a process of upgrading and optimizing model training. First, when this function runs, the software at the edge and central ends has been deployed and debugged, and the business function is running normally. Only the model accuracy needs to be continuously upgraded and iterated.
[0060] The following takes the iterative training and optimization of the fireworks detection model as an example for illustration.
[0061] First, when the fireworks models deployed on each terminal are working, the models will detect the video images transmitted by the edge scene cameras frame by frame. When the models detect that there is a fireworks occurrence, the platform software saves the alarm pictures and sends out alarm messages, and at the same time sends the alarm messages (pictures and detection information, etc.) to the central end platform.
[0062] When the central end receives the edge alarm information, it will add the alarm pictures to the corresponding sample picture library according to the alarm type, such as the fireworks detection sample picture library. At the same time, it generates a corresponding txt file (text file, plain text file), and adds the corresponding detection labels and coordinate box information. The alarm pictures sent by each edge terminal received by the central end are all processed in the same way.
[0063] At the same time, the training module located at the central end will use the picture understanding function of the multi-modal pre-trained language model to verify the pictures and detection information added to the sample picture library to replace the manual verification process. First, obtain a picture of fireworks detection and related detection information from the sample picture library, such as labels and detection box coordinates, etc. Then, crop the image within the picture box according to the coordinate box information and place the cropped part at the original coordinate box position of a white background picture of the same size as the original picture to form a new picture. Then generate the pre-trained language model question: Question1 = "Are there any fireworks in the picture? The format of output should be like {“yes”, “no”}”, which is used to detect whether there is a specific object, such as fireworks, in the image, and requires the pre-trained language model to simply answer whether the specified object is detected in the picture. The output format is a simple binary form: “yes” or “no”, which helps to quickly determine whether the picture contains the target object.
[0064] Question2 = "What object is in the picture? The format of output should be like { “label”: the name of this object}”. It aims to identify and return the name of the main object present in the picture. Through Question2, the object label in the picture can be obtained (for example, “fireworks”, “helmet”, etc.), confirming the main content of the picture.
[0065] Question3 = "Detect all objects in the image and return their locations in the form of coordinates. The format of output should be like {“bbox”: [x, y, w, h], “label”: the name of this object}". It is used to comprehensively detect all objects in the picture and requires the pre-trained language model to provide the location information (bounding box coordinates) of each detected object and its corresponding label. The detailed output can help precisely understand the locations and categories of each object in the picture, providing more detailed information than the previous two questions and supporting further data analysis and processing.
[0066] Among them, in Question1, “fireworks” is replaced according to different models. For example, in helmet detection, it is changed to “helmet”.
[0067] Finally, taking the newly generated picture and the questions of the pre-trained language model as parameters, call the interface of the locally deployed pre-trained language model, and analyze the relevant detected objects and coordinate information returned by the pre-trained language model.
[0068] Based on the returned result information, if the return result is not empty and the detection result returned by the pre-trained language model is the same as the detection result uploaded by the visual model, calculate the new coordinates according to the following formula (1): , (1) Among them, is used to calculate the new coordinates of the detection box, and are weight coefficients; are the coordinates returned by the pre-trained language model; are the detection coordinates uploaded by the visual model.
[0069] After the calculation is completed, update the information of the corresponding picture's txt file using the new coordinate information. If the return result is not empty and the target types are inconsistent, delete the txt coordinate box information and the detection category of the corresponding picture. If the return result of the pre-trained language model is empty, also delete the txt coordinate box information and the detection category of the corresponding picture. Finally, after verifying all the detection boxes in a picture, if there is no coordinate box information in the txt file, add the picture to the negative sample library.
[0070] When the increased number of samples in the training sample library meets the threshold requirement, enter the next sample preparation stage. At this time, check the number of samples in the negative sample library. If the number of negative samples is less than the set value, such as 1 / 5 of the sample number, perform negative sample enhancement. Randomly extract the difference in the number of pictures from the sample library. Based on these pictures, use the picture generation function of the pre-trained language model to reconstruct them, delete the existing target objects in the pictures, and only retain the background picture information as negative sample materials. At the same time, for the scenarios where the model (such as fireworks recognition) is most likely to misidentify, such as the background of sunlight during the day and car headlights at night, generate picture text descriptions, and use the picture generation function of the pre-trained language model to generate negative sample pictures.
[0071] Once the sample preparation is complete, call the training script to train the model. After the model training is completed, perform a preliminary verification. After the verification is completed, perform model conversion, and then perform a one-key batch distribution to complete this round of iterative upgrade.
[0072] The system flow chart of this method is as Figure 2 shown.
[0073] Specifically, the main steps of sample picture upload, collection, verification, training, conversion, etc. are as follows: Step 1: Deploy each edge-side algorithm module and run the algorithm model for target recognition and detection (such as fireworks recognition).
[0074] Step 2: The edge-side model sends the recognized target alarm picture samples (such as fireworks) and detection information to the central end.
[0075] Step 3: When the central end receives the alarm picture information, perform preprocessing on the sample picture information extraction to obtain the picture, label, and coordinate box information, and generate a txt file at the same time.
[0076] Step 4: Add the picture and the corresponding detection information to the sample picture library.
[0077] Step 5: Add the picture and the corresponding detection information to the pre-trained language model verification list.
[0078] Step 6: Obtain the picture samples and detection information one by one from the pre-trained language model verification list for verification.
[0079] Step 7: Update the training sample image library and the corresponding txt file information according to the verification results.
[0080] Step 8: Check the number of newly added training samples. When the number of newly added training samples is greater than the set threshold, go to Step 9 to prepare the training samples. Otherwise, continue to execute this step.
[0081] Step 9: Prepare the training samples, that is, perform negative sample enhancement to meet the training requirements.
[0082] Step 10: Call the model training script to train the model.
[0083] Step 11: After the verification is completed, perform model conversion to convert it into a model format suitable for the edge acceleration card.
[0084] Step 12: One-click model distribution, batch distribute the latest model to each edge device.
[0085] Step 13: The edge device receives the model and performs verification.
[0086] Step 14: If the verification is qualified, perform model replacement, and thus a round of iterative upgrade of the model is completed.
[0087] As Figure 3 shown, the main steps for verifying the multi-modal pre-trained language model are as follows: S201. Obtain the current image and the current detection information (label and coordinate box information) from the pre-trained language model verification list.
[0088] S202. Crop the image within the coordinate box according to the coordinate box information and place the cropped part at the original coordinate box position of a white background image of the same size as the original image to generate a new image.
[0089] S203. Generate model verification questions: Question1, Question2, and Question3.
[0090] S204. Use the newly generated image and questions to call the multi-modal pre-trained language model interface.
[0091] S205. Analyze the results feedback by the pre-trained language model.
[0092] S206. If the feedback target result is empty, delete the coordinate information in the corresponding txt file. If the feedback target result is not empty, go to S207.
[0093] St207. Extract the target type and coordinate information, and compare it with the upload result of the visual model. If the target types are the same, go to S210. If the types are different, go to 208 and delete the coordinate information in the corresponding txt file.
[0094] S208. Delete the coordinate information corresponding to the txt file.
[0095] S209. Determine whether all detection boxes in this figure have been verified. If they have been verified, proceed to step S213. If not, go to S202.
[0096] S210. Calculate the final coordinate box. At this time, there are two sets of coordinate box information. Take the weight of the coordinate box of the visual model as 0.25 and the weight of the coordinate box information of the pre-trained language model as 0.75, and calculate the final detection target coordinate box by weighted calculation.
[0097] S212. Determine whether all detection boxes in this figure have been verified. If they have been verified, proceed to step S213. If not, go to S202.
[0098] S213. If the txt file corresponding to this figure has no coordinate pivot information, add this picture to the negative sample library.
[0099] As Figure 4 shown, the main steps of negative sample enhancement of the modal pre-trained language model are as follows: S301. Obtain the difference between the number of negative samples and the set value.
[0100] S302. Determine whether the difference is less than 1 / 3 of the set value. If it is less, go to S304; otherwise, go to S303.
[0101] S303. Randomly copy the above-mentioned number of sample pictures from the sample library.
[0102] S304. Set the difference quantity to 1 / 3 of the set total.
[0103] S305. Use the copied pictures as the background, call the picture generation function of the pre-trained language model, and generate pictures without recognition targets.
[0104] S306. Import the key environmental background information that the current model is most likely to misjudge.
[0105] S307. Use the imported key background information that the current model is most likely to misjudge to generate picture description information. Use the picture description information and use the pre-trained language model to generate negative sample pictures with a quantity of 1 / 2 of the difference.
[0106] S308. Determine whether the number of negative samples is less than 1 / 5 of the total training samples. If it is satisfied, end the negative sample enhancement; otherwise, go to S309.
[0107] S309. Randomly delete some negative sample pictures so that the total number is less than or equal to 1 / 5 of the total training samples.
[0108] The above solution is applied to the cluster version of the edge intelligence management platform to realize automatic training and upgrading of AI vision models, including automatic sample push, sample classification, sample verification, negative sample enhancement, model training, model conversion, batch model distribution and upgrading, etc., to meet the integrated push and training requirements of the edge management platform.
[0109] The central / cluster version of the Yuanzhi management platform is generally deployed on the server in the central computer room, while the stand-alone version is deployed on device terminals scattered around. When the algorithm model deployed on the edge starts running, the real-time detected target samples will be pushed to the cluster version, i.e. the central end. A series of operations such as image sample classification, sample verification, negative sample enhancement, model training, and model conversion will be completed at the central end. Finally, it will be sent back to each edge end to realize the iterative upgrade of the model. This process does not require human intervention, thus solving the problem of manual collection, labeling, and verification of sample images, greatly improving development efficiency, and saving manpower and material resources.
[0110] This process uses the automatic target detection samples of the algorithms configured on each edge to collect samples, thus eliminating the need for manual sample collection. The more devices deployed on the edge, the more samples can be collected, which can achieve sampling diversity and greatly improve the model generalization capability.
[0111] In traditional model training, sample labeling and verification are the most time-consuming and labor-intensive. This embodiment realizes preliminary labeling of samples through the first step of sample collection. At the same time, for the incorrectly labeled parts, it reasonably constructs question and answer styles and combines the image understanding capabilities of the multimodal pre-trained language model to realize the verification of sample images, and makes modifications to the parts with preliminary verification errors to replace the manual verification process.
[0112] As for the situation where negative samples are scarce, one approach is to use some sample images as background colors and call the image generation function of the pre-trained language model to generate images without recognition targets. At the same time, the imported current model is most likely to misjudge key background information to generate image descriptions, and the pre-trained language model is used to generate negative sample images that meet the background requirements, thereby enriching the number of negative sample training and greatly improving the accuracy of the final model.
[0113] Through automatic sample collection and labeling, automatic verification, automatic negative sample enhancement, and then calling the training script to implement model training and conversion, and finally one-click distribution, the purpose of automatic iteration and upgrading can be achieved.
[0114] The embodiment of the present invention also provides a visual model upgrading device, such as Figure 5 As shown, the visual model upgrading device 10 includes: an acquisition module 401 , a generation module 402 and an upgrading module 403 .
[0115] Among them, the acquisition module 401 is used to acquire the picture sample and the picture detection result uploaded by the target terminal; the generation module 402 is used to generate a training sample according to the picture sample and the picture detection result, update the database according to the training sample, and identify the update time of the visual model; the upgrade module 403 acquires the number of newly added training samples in the database since the update time. When the number of training samples is greater than the number threshold, the visual model is trained using the training samples in the database, and the visual model of the target terminal is upgraded using the trained visual model.
[0116] In an embodiment of the present invention, the generation module 402 is further configured to: acquire the picture sample and the picture detection result in the training sample; extract the coordinate box information in the picture detection result, and perform cropping processing on the picture sample using the coordinate box information to obtain a new picture; generate model verification questions according to the type of the visual model, and call the pre-trained language model on the server according to the model verification questions and the new picture; verify the training sample according to the feedback result of the pre-trained language model; update the database using the verification result.
[0117] In an embodiment of the present invention, the generation module 402 is further configured to: determine the picture frame of the picture sample according to the coordinate box information; crop the image in the area of the picture frame, and place the area image on the picture template to obtain a new picture, where the position of the area image on the picture sample and the picture template is the same.
[0118] In an embodiment of the present invention, the generation module 402 is further configured to: identify whether the feedback result is empty; if the feedback result is empty, delete the coordinate box information of the corresponding picture sample; if the feedback result is not empty, compare whether the label information in the picture detection result and the feedback result is consistent. If the label information is inconsistent, delete the coordinate box information of the corresponding picture sample. If the label information is consistent, calculate the new coordinate box information of the picture sample according to the coordinate box information in the picture detection result and the feedback result.
[0119] In an embodiment of the present invention, the generation module 402 is further configured to: acquire the respective weights of the visual model and the pre-trained language model; calculate the new coordinate box information of the picture sample by weighted calculation according to the coordinate box information in the picture detection result, the coordinate box information in the feedback result, and the respective weights of the visual model and the pre-trained language model.
[0120] In an embodiment of the present invention, the generation module 402 is further configured to: use the picture sample with the coordinate box information deleted as a negative sample; use the picture sample carrying the new coordinate box information and the label information as a positive sample; add the negative sample and the positive sample to the database.
[0121] In an embodiment of the present invention, the upgrade module 403 is further configured to: select positive samples and negative samples from a database; generate a training data set according to the positive samples and negative samples; and train a visual model by using the training data set.
[0122] In an embodiment of the present invention, an enhancement module is further included, and the enhancement is further configured to: before training the visual model by using the training data set, based on the training samples in the database, enhance the number of negative samples in the training data set.
[0123] In an embodiment of the present invention, the enhancement module is further configured to: obtain the number of negative samples in the training data set; if the number of negative samples is less than a target number, calculate the difference between the number of negative samples and the target number; randomly select training samples from the database based on the difference, and enhance the number of negative samples in the training data set based on the training samples.
[0124] In an embodiment of the present invention, the enhancement module is further configured to: if the difference is less than a preset number, randomly select the number of training samples equal to the difference from the database; if the difference is greater than or equal to the preset number, randomly select the preset number of training samples from the database.
[0125] In an embodiment of the present invention, the enhancement module is further configured to: extract sample pictures from the training samples; input the sample pictures into a pre-trained language model of the server, and use the picture generation function of the pre-trained language model to generate pictures without label information; and generate negative samples according to the picture description information and the pictures without label information.
[0126] In an embodiment of the present invention, the enhancement module is further configured to: obtain misjudgment data of the visual model; and generate picture description information according to the misjudgment data.
[0127] In an embodiment of the present invention, the enhancement module is further configured to: after enhancing the number of negative samples in the training data set based on the training samples, obtain the total number of training samples in the training data set; set an enhancement stop threshold according to the total number of training samples, where the enhancement stop threshold is less than the total number of training samples; if the number of negative samples is less than the enhancement stop threshold, stop enhancing the number of negative samples in the training data set, and if the number of negative samples is greater than or equal to the enhancement stop threshold, delete some negative samples in the training data set.
[0128] For the description of the features in the corresponding embodiment of the visual model upgrade device, reference may be made to the relevant description in the corresponding embodiment of the visual model upgrade method, which will not be elaborated here one by one.
[0129] An embodiment of the present invention further provides an electronic device, such as Figure 6As shown, it includes a memory 501 and a processor 502. A computer program is stored in the memory 501, and the processor 502 is configured to run the computer program to execute the steps in the above-described embodiments of the visual model upgrade method.
[0130] An embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described embodiments of the visual model upgrade method when running.
[0131] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs and other various media that can store computer programs.
[0132] An embodiment of the present invention also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the visual model upgrade method.
[0133] An embodiment of the present invention also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the visual model upgrade method.
[0134] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0135] The above has introduced in detail a visual model upgrade method, an electronic device, a storage medium, and a program product provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A method for upgrading a visual model, characterized in that, The visual model is deployed on the target terminal, and the visual model outputs the image detection results of the image samples. The method is applied to the server, where the method includes: Obtain the image samples and image detection results uploaded by the target terminal; Generate training samples according to the image samples and the image detection results, update the database according to the training samples, and identify the update time of the visual model; Obtain the number of training samples newly added to the database since the update time. When the number of training samples is greater than the quantity threshold, use the training samples in the database to train the visual model, and use the trained visual model to upgrade the visual model of the target terminal.
2. The visual model upgrade method according to claim 1, wherein The updating the database according to the training samples further includes: Obtain the image samples and image detection results in the training samples; Extract the coordinate box information in the image detection results, and perform cropping processing on the image samples using the coordinate box information to obtain new images; Generate model verification questions according to the type of the visual model, and call the pre-trained language model on the server according to the model verification questions and the new images; Verify the training samples according to the feedback results of the pre-trained language model; Update the database using the verification results.
3. The visual model upgrade method according to claim 2, wherein The performing cropping processing on the image samples using the coordinate box information to obtain new images includes: Determine the image box of the image sample according to the coordinate box information; Crop the image within the image box, and place the cropped image on an image template to obtain a new image, where the position of the cropped image on the image sample and the image template is the same.
4. The visual model upgrade method according to claim 2, wherein The verifying the training samples according to the feedback results of the target pre-trained language model includes: Identify whether the feedback result is empty; If the feedback result is empty, delete the coordinate box information of the corresponding image sample; If the feedback result is not empty, compare whether the label information in the image detection result and the feedback result is consistent. If the label information is inconsistent, delete the coordinate box information of the corresponding image sample. If the label information is consistent, calculate the new coordinate box information of the image sample according to the coordinate box information in the image detection result and the feedback result.
5. The visual model upgrade method according to claim 4, wherein The calculating the new coordinate box information of the image sample according to the coordinate box information in the image detection result and the feedback result includes: Obtain the respective weights of the visual model and the pre-trained language model; Weightedly calculate the new coordinate box information of the image sample according to the coordinate box information in the image detection result, the coordinate box information in the feedback result, and the respective weights of the visual model and the pre-trained language model.
6. The visual model upgrade method according to claim 2, wherein The updating the database using the verification results includes: Use the image samples with the coordinate box information deleted as negative samples; Use the image samples with the new coordinate box information and label information as positive samples; Add the negative samples and the positive samples to the database.
7. The visual model upgrade method according to claim 1, characterized in that The training samples include positive samples and negative samples. The training the visual model using the training samples in the database includes: Select positive samples and negative samples from the said database; Generate a training data set based on the said positive samples and negative samples; Use the said training data set to train the said visual model.
8. The visual model upgrade method according to claim 7, wherein Before using the said training data set to train the said visual model, it further includes: Enhance the number of negative samples in the said training data set based on the training samples in the said database.
9. The visual model upgrade method according to claim 8, wherein The enhancing the number of negative samples in the said training data set based on the training samples in the said database includes: Obtain the number of negative samples in the said training data set; If the number of negative samples is less than the target number, calculate the difference between the number of negative samples and the target number; Randomly select training samples from the said database based on the said difference, and enhance the number of negative samples in the said training data set based on the said training samples.
10. The visual model upgrade method according to claim 9, wherein, The randomly selecting training samples from the said database based on the said difference includes: If the said difference is less than the preset number, randomly select the said preset number of training samples from the said database; If the said difference is greater than or equal to the said preset number, randomly select the said difference number of training samples from the said database.
11. The visual model upgrade method according to claim 9, wherein, The enhancing the number of negative samples in the said training data set based on the said training samples includes: Extract the sample pictures in the said training samples; Input the said sample pictures into the pre-trained language model of the said server, and use the picture generation function of the said pre-trained language model to generate pictures without label information; Generate negative samples according to the picture description information and the pictures without label information.
12. The visual model upgrade method according to claim 11, wherein Before generating negative samples according to the picture description information and the pictures without label information, it further includes: Obtain the misjudgment data of the said visual model; Generate the said picture description information according to the said misjudgment data.
13. An electronic device, characterized in that, It includes: A memory for storing computer programs; A processor for implementing the steps of the visual model upgrade method as described in any one of claims 1 to 12 when executing the said computer programs.
14. A computer-readable storage medium, characterized in that, A computer program is stored in the said computer-readable storage medium, wherein the said computer program implements the steps of the visual model upgrade method as described in any one of claims 1 to 12 when executed by a processor.
15. A computer program product, comprising a computer program, characterized in that, The said computer program implements the steps of the visual model upgrade method as described in any one of claims 1 to 12 when executed by a processor.
Citation Information
Patent Citations
Target detection model training method and device, equipment and storage medium
CN117710994A
Multi-modal large model processing method and device, storage medium and program product
CN119314117A
Updating method, device and equipment of coal mine underground AI visual model and medium
CN119540728A
Visual training system and program
JP2007068579A
Apparatus, method, and computer program for processing visual event data
WO2024200170A1