Method and device, apparatus, vehicle and medium for generating a classification model
The method addresses inefficiencies in vehicle object detection by using a semantic text-image matching model to generate a classification model for specific scenarios, enhancing accuracy and efficiency in vehicle object detection.
Patent Information
- Application Number
- DE102024211705
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-12-09
- Publication Date
- 2025-06-26
AI Technical Summary
Current data mining solutions for vehicle object detection are inefficient due to manual data search and labeling, which is time-consuming, labor-intensive, and prone to inconsistencies and false identifications, leading to insufficient accuracy and generalization.
A method and apparatus for generating a classification model by capturing images associated with a target text using a semantic text-image matching model, generating an image pattern set with category labels, and creating a classification model for the target scenario based on a pre-trained model using linear recognition.
This approach enables rapid generation of customized classification models for specific scenarios, improving flexibility and accuracy while ensuring the efficiency and stability of the model generation process.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
FIELD OF THE INVENTIONThe embodiments of the present disclosure generally relate to the field of computers, and more specifically to a method and apparatus, apparatus, vehicle, and medium for generating a classification model.PRIOR ARTObject detection is a computer vision technology that allows vehicles, pedestrians, traffic signs, etc. to be localized and identified in visual information such as images or videos. Such technologies may help vehicle driving systems (e.g., advanced driver assistance systems (ADAS) and automated driving systems (ADAS)) better understand the driving environment, thereby making safer and more efficient decisions, and taking appropriate actions.Object detection in vehicles is an important part of assisted driving and autonomous driving techniques. It uses techniques such as image recognition and machine learning to recognize and identify important objects during the driving process and thus provide important decision bases for the driving system. With the continued advancement of computer vision and machine learning technologies, the accuracy and real-time performance of vehicle object detection has also been constantly improved.DISCLOSURE OF THE INVENTIONThe embodiments of the present disclosure provide a method, apparatus, apparatus, and medium for generating a classification model.In a first aspect of the present disclosure, a method of generating a classification model is provided. The method includes capturing a plurality of images associated with a target text using a semantic text-image matching model, the target text indicating a target scenario. The method further includes generating an image pattern set including image patterns with category labels by determining the category to which each image of the plurality of images belongs. The method further includes generating a classification model for the target scenario based on the image pattern set, wherein generating the classification model is based on a pre-trained model using linear recognition.According to a second aspect of the present disclosure, an apparatus for generating a classification model is provided. The apparatus includes an image acquisition module configured to acquire multiple images associated with a target text using a semantic text-image matching model, the target text indicating a target scenario. The apparatus further comprises a pattern set generation module configured to generate an image pattern set comprising image patterns with category labels by determining the category to which each image of the plurality of images belongs. The apparatus further comprises a model generation module configured to generate a classification model for the target scenario using the image pattern set, wherein the classification model is a pre-trained model using linear recognition.In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes at least one processor. The electronic device further comprises a memory coupled to the at least one processor and having instructions stored therein, which instructions, when executed by the at least one processor, cause the device to perform the steps of the method of the first aspect of the present disclosure.In a fourth aspect of the present disclosure, a vehicle is provided. This vehicle includes an electronic device of the third aspect of the present disclosure.According to a fifth aspect of the present disclosure, there is provided a computer readable storage medium. Stored on this computer readable storage medium are computer executable instructions, which computer executable instructions are executed by a processor to implement the steps of the method of the first aspect of the present disclosure.DESCRIPTION OF THE FIGURESBy describing the exemplary embodiments of the present disclosure in more detail in conjunction with the figures, the above and other objects, features, and advantages of the present disclosure will become more apparent, wherein like or similar reference numerals generally represent like or similarly shaped components, components, and the like in the exemplary embodiments of the present disclosure. FIG. 1 shows a schematic representation of an example environment in which methods and / or apparatuses of embodiments of the present disclosure may be implemented; FIG. 2 is a flowchart of an apparatus for generating a classification model according to an embodiment of the present disclosure; FIG. 3 is a diagram illustrating a process for generating a classification model according to an embodiment of the present disclosure; FIG. 4 is a diagram illustrating a search operation for an image set according to an embodiment of the present disclosure; FIG. 5 illustrates the generation process of an image pattern set according to an embodiment of the present disclosure; FIG. 6 illustrates the training process of a classification model according to an embodiment of the present disclosure; FIG. 7 illustrates the test process of a classification model according to an embodiment of the present disclosure; FIG. 8 illustrates the prediction process of a classification model according to an embodiment of the present disclosure; FIG. 9 is a schematic diagram of an apparatus for generating a classification model according to an embodiment of the present disclosure; and FIG. 10 shows a schematic block diagram of an example apparatus suitable for use in implementing the embodiments of the present disclosure.In the figures, like or corresponding reference numerals designate the same or corresponding parts.DETAILED DESCRIPTION OF THE EMBODIMENTSHereinafter, the embodiments of the present disclosure will be described in more detail with reference to the figures. Although certain embodiments of the present disclosure are illustrated in the figures, it should be understood that the present disclosure may be practiced in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a better and complete understanding of the present disclosure. It should be understood that the figures and embodiments of the present disclosure are merely examples and are not intended to limit the scope of the disclosure.In describing the embodiments of the present disclosure, the term "comprise" and its variations are to be understood as open-ended inclusion, i.e., "comprising, but not limited to.". The term "based on" should be understood to mean "at least partially based on.". The terms "an exemplary embodiment" or "this exemplary embodiment" are to be understood as "at least one exemplary embodiment". The terms "first / r / s", "second / r / s", etc. may represent different or identical objects, unless expressly stated otherwise.As described above, the object detection function in vehicles helps improve performance of ADAS and autonomous AD driving systems. It uses techniques such as image recognition and machine learning to identify important objects in the driving environment (such as vehicles, pedestrians, traffic signs, etc.), thereby better understanding the driving environment, which causes the corresponding ADAS or AD system to make more secure and effective decisions and take appropriate action.Object detection for vehicles may be based on neural networks or other machine learning models (such as convolutional neural networks, backpropagation neural networks, etc.), where these models are capable of extracting useful features from visual information such as images or videos and using these extracted features to identify important objects. Large amounts of high-quality data are required for the creation of high-performance models, and it is therefore necessary to use a suitable data mining strategy in order to obtain as many valuable data as possible, which is of decisive importance both for the accuracy and for the generalizability of the model.Although the provision of data mining strategies for vehicle object detection has been continuously advanced in recent years as an indispensable part of the functional development of ADAS or AD systems, current data mining solutions still have some problems that result in not completely meeting the requirements of vehicle object detection, such as insufficient accuracy, insufficient generalization, etc.Prior art solutions manually search the raw data and manually label the searched data to form a pattern set, and then use the manually generated pattern set to train a model. Such methods for searching raw data and creating sample sets may not be able to efficiently process numerous and varied data. Generally, the amount of data required to detect vehicle objects is very large, and the data includes a variety of different scenarios (e.g., tunnels, freeways, etc.) and conditions (e.g., weather conditions, traffic jams, etc.). The manual search and labeling is slow, time and labor intensive and cannot ensure consistency and stability. Moreover, the exclusive resort to manual work inevitably leads to false identification and forgetful patterns. Moreover, conventional modeling requires complete matching of most or even all of the parameters of the model, resulting in a tedious process that is detrimental to accuracy and robustness.To solve at least the above and other potential problems, the embodiments of the present disclosure provide a solution for generating a classification model. The solution for generating a classification model according to the embodiments of the present disclosure comprises acquiring a plurality of images in connection with a target text by means of a semantic text-image matching model, wherein the target text indicates a target scenario. The solution further comprises generating an image pattern set comprising image patterns with category labels by determining the category to which each image of the plurality of images belongs. The solution further includes generating a classification model for the target scenario based on the image pattern set, wherein generating the classification model is based on a pre-trained model using linear recognition. In this way, it is possible to rapidly generate customized classification models for particular or concrete scenarios that meet the actual requirements, thereby improving flexibility and accuracy while ensuring the efficiency and stability of the model generation process.FIG. 1 shows a schematic representation of an example environment 100 in which methods and / or processes of embodiments of the present disclosure may be implemented. As shown in FIG. 1, the example environment 100 may include a vehicle 110, an image set 120, a computing device 130, and a storage device 140, and these components may be coupled together for interaction as shown in FIG. 1. It should be understood that only limited components are shown in the example environment 100 for implementing the embodiments of the present disclosure to facilitate understanding and illustration, and that the embodiments of the present disclosure are not limited thereto. For example, the example environment 100 may further include a display device (not shown) configured to display classification results.According to an embodiment of the present disclosure, the vehicle 110 may be any type of motorized or non-motorized transport means capable of transporting persons and / or objects and movable. The vehicle 110 typically includes one or more wheels, one or more seats, one or more load bearing structures (e.g., a car, cabin, etc.), one or more propulsion systems (e.g., an engine, an electric motor, etc.), one or more control systems (e.g., a steering wheel, an accelerator pedal, etc.), and one or more safety systems (e.g., seatbelts, airbags), etc.As shown in FIG. 1, the vehicle 110 is an automobile. However, this is illustrative only and is not limiting. The vehicles 110 may include, but are not limited to, passenger cars, trucks, off-road vehicles, sport utility vehicles, motor cycles, and the like. Moreover, the vehicle 110 may be based on fossil energy or clean energy, or a combination thereof. Fossil-based vehicles are mainly understood to mean vehicles that use fossil fuels such as oil and natural gas as energy sources, such as conventional gasoline vehicles, diesel vehicles, etc. Clean-energy-based vehicles refer to vehicles that use clean energy as a drive source, such as electric vehicles, hydrogen fuel cell vehicles, solar vehicles, etc.According to the embodiments of the present disclosure, the vehicle 110 may include multiple sensors for implementing driving specific functions such as autonomous driving or assisted driving. These sensors may cooperate with the computing device 130 to enable precise sensing of the environment surrounding the vehicle 110, decision control, and the like. The plurality of sensors included in the vehicle 110 may include vision sensors (e.g., cameras, infrared sensors, depth sensors, etc.) for capturing an image set 120 that indicates the driving environment and scenarios of the vehicle 110 during the trip.According to an exemplary embodiment of the present disclosure, the image set 120 may include images indicating various different scenarios and conditions during travel of the vehicle 110, and the like. Examples of such scenarios may include, but are not limited to, tunnels, intersections, transits, etc. The image set 120 may be stored in the storage device 140 and retrieved via the computing device 130. In some embodiments, the image set 120 may be presented in the form of a video comprising image frames. As discussed above, the images in the image set 120 may be captured by visual sensors on the vehicle 110 such as cameras, infrared sensors, depth sensors, etc., whereby images in the image set 120 may include red green blue (RGB) images, infrared images, and depth images. It should be understood that this is not limiting and the images in the image set 120 may also be obtained in other ways, for example, by a data augmentation process. Moreover, the size and format of the image are not limited in the present embodiments, so that it is sufficient to select the appropriate size and format according to the actual need.According to embodiments of the present disclosure, computing device 130 may include an in-vehicle computing device, an out-of-vehicle computing device, and combinations thereof. The computing device 130 may have computational capabilities and be configured to perform classification model generation according to embodiments of the present disclosure. During execution of classification model generation according to embodiments of the present disclosure, computing device 130 may access memory device 140 and perform corresponding computations. It should be understood that computing device 130 is shown in FIG. 1 as a computing device, but this is illustrative only and is not limiting, and a larger number of computing devices may also be present in example environment 100. In the following, respective operations of the computing device 130 will be described in more detail.In an exemplary and non-limiting manner, the computing device 130 may be, but is not limited to, a personal computer, a laptop computer, a server computer, a mobile device (e.g., a smartphone, a tablet computer, etc.), a portable electronic device, a multimedia player, a personal digital assistant (PDA), a smart home device, an consumer electronic product, or a distributed computing environment comprising one or more of the aforementioned devices.According to the embodiments of the present disclosure, the storage device 140 may be configured to store an image set 120 of the vehicle 110, the models to be used by the computing device 130, and their parameters and classification results. It should be understood that the storage device 140 is depicted in FIG. 1 as a storage device, but this is for illustrative purposes only and is not limiting, and a larger number of storage devices may be present in the example environment 100.In an exemplary and non-limiting manner, the storage device 140 may include, but is not limited to, local storage devices, remote storage devices, and combinations thereof. In some embodiments, the plurality of storage devices in the storage device 140 may include, but are not limited to, mechanical hard disks (HDDs), solid state hard disks (SSDs), etc., and some of the plurality of storage devices may be locally located while others may be remotely located, being interconnected, e.g., via wires or networks.Above, with reference to FIG. 1, example environment 100 in which methods and / or processes of embodiments of the present disclosure may be implemented has been described. A flowchart of a method 200 for generating a classification model according to an exemplary embodiment of the present disclosure is described below with reference to FIG. 2. The method 200 can effectively take into account specific or concrete scenarios in the generation of a classification model, thereby increasing the flexibility and accuracy of the model and thereby improving the quality of the data mining. Additionally, the classification model is based on incremental training to ensure the efficiency and stability of the model generation process.In block 210, a plurality of images are acquired in connection with a target text using a semantic text-image matching model, wherein the target text specifies a target scenario. According to the embodiments of the present disclosure, text associated with a target scenario (hereinafter also referred to as target text) can be obtained. In some embodiments, such text may be received from a user, for example, via input devices such as keyboards and microphones. After receiving the target text, the computing device 130 may analyze the target text and, based thereon, retrieve a plurality of corresponding images from the image set 120 using a semantic text-image matching model.According to the embodiments of the present disclosure, by the semantic text-image matching model, the text semantics can be extracted from the target text, and the semantic level matching and the matching between the target text and the images in the image set can be performed, thereby finding images associated with the target text. With such text-to-image semantic matching, the semantic relationship between text and image can be better understood, and accurate image feedback can be given. The text-to-image semantic matching method according to the embodiments of the present disclosure will be described in detail below.In block 220, generating an image pattern set comprising image patterns with category labels is performed by determining the category to which each image of the plurality of images belongs. According to embodiments of the present disclosure, labeling processing of the plurality of images associated with the target text obtained in 210 may be performed so that each of these images is automatically assigned a corresponding category label, wherein such labeling process may be based on computer vision technology for images. Next, the generation process of an image pattern set according to an embodiment of the present disclosure will be described in more detail.In block 230, a classification model for the target scenario is generated based on the image pattern set, wherein the generation of the classification model is based on a pre-trained model using linear recognition. According to embodiments of the present disclosure, based on the image pattern set including image patterns with category labels generated in 220, the parameter set of the classification model may be adjusted, which in turn generates a classification model for a specific scenario, wherein the generation of this classification model is based on a pre-trained model using linear recognition (linear probing). In other words, the classification model is already trained before the generation.Linear detection is a technique for model fine tuning. In linear recognition, parameters may be optimized for particular tasks based on a pre-trained model, such an adaptation process involving only a portion of the model. According to the embodiments of the present disclosure, in the linear recognition based on a classification model pre-trained with a large amount of data, the model parameters for the target scenario indicated by the target text may be finely tuned, which fine tuning affects only a small number of layers of the model or a small number of parameters in a parameter set. In this way, the model's pretraining results may be extensively utilized and retraining of the entire model may be avoided, thereby saving computational resources and time. At the same time, by fine tuning a small number of layers or parameters of the model, we can improve the adaptability of the model to specific scenarios, while preserving the stability of the model.By the method for generating a classification model according to the embodiments of the present disclosure, it is possible to automatically form a pattern set of correspondingly labeled patterns for supervised or semi-supervised learning for certain scenario requests indicated by text while maintaining model stability, perform model adaptation for certain scenario requests, and quickly provide classification models with improved flexibility and accuracy. Hereinafter, the individual embodiments of the classification generation process according to the embodiments of the present disclosure will be described in more detail.FIG. 3 is a diagram of a process for generating a classification model 300 according to an embodiment of the present disclosure. As exemplarily shown in FIG. 3, the classification model generation process 300 may include a search sub-process 310, a dataset generation sub-process 320, a training sub-process 330, a test sub-process 340, and a prediction sub-process 350. The classification model generation process 300 and its sub-processes may be abstracted into classification model generation units and corresponding sub-units for the sub-processes (e.g., search sub-unit, dataset generation sub-unit, etc.), which may be software-based implemented components or systems for generating classification models, and which may be executed on a device with computational capabilities (such as the computing device 130).According to embodiments of the present disclosure, in the search sub-process 310, a text description of the scenario of interest may be input, and multiple images corresponding to the scenario of interest in the image set 120 may be searched based on the text description, where the image set 120 may include images captured by the vehicle 110 visual sensors as described above, but also, for example, original images extracted from sequence data, and the like. It should be understood that the embodiments of the present disclosure do not limit obtaining images in the image set 120. The multi-modal search sub-process 310 can use the text information to roughly filter the image set for the scenario of interest. The search sub-process 310 according to the embodiments of the present disclosure will be described in more detail below.According to an embodiment of the present disclosure, in the dataset generation sub-process 320, each image among the multiple images corresponding to the target scenario found at 310 may be automatically assigned a corresponding category label. For example, for a tunnel input scenario, each image may be assigned the designation "input" or "non-input", or for a tunnel environment scenario, each image may be assigned the designation "input", "output", "interior", or "other". In some embodiments, such automatic tagging processing may be folder-based. It should be understood that the examples of scenarios and categories described herein are for convenience and illustration only and are not to be considered limiting. The dataset generation sub-process 320 may generate a sample set of samples with category labels for subsequent training and testing. Hereinafter, the record generation sub-process 320 according to an embodiment of the present disclosure will be described in more detail.According to an embodiment of the present disclosure, in the training sub-process 330, a portion of the image pattern set generated in 320 may be used to train the classification model. Because such image patterns relate to the target scenario, the trained classification model may be adapted to the target scenario by this training sub-process 330 and have a higher degree of adjustment to the target scenario, resulting in improved accuracy. In other words, the model is pre-trained using large amounts of data, and thereafter fine tuning of the parameters is performed during training for the target scenario indicated by the target text. The training sub-process 330 may better exploit the results of pretraining the model, and only a few layers or parameters of the model are fine tuned, which saves computing resources and time while maintaining model stability while maintaining high accuracy. The training sub-process 330 according to an embodiment of the present disclosure will be described in more detail below.According to an embodiment of the present disclosure, in test sub-process 340, another portion of the image pattern set generated in 320 may be used to test the classification model trained in 330 and evaluate the training effect. In some embodiments, an accuracy threshold may be predefined and the performance of the trained classification model may be checked based on the predefined accuracy threshold. For example, if at 345 the model accuracy satisfies the expectation indicated by the predefined accuracy threshold, the classification model is used for subsequent prediction processes. On the other hand, if the model accuracy does not meet the expectation, the classification model must be further adjusted until it meets the expectation. In some embodiments, in response to model accuracy being less than expected, classification model generation process 300 returns to training sub-process 330 to perform retraining. The test sub-process 340 according to an embodiment of the present disclosure will be described in more detail below.According to an embodiment of the present disclosure, in prediction sub-process 350, prediction may be performed using the trained classification model verified by test sub-process 340 to determine the category to which the input image belongs. In some embodiments, input images associated with a target scenario may be received and input to a trained classification model to predict the category to which each input image under the target scenario belongs. In this way, the category to which each input image belongs can be determined by the generated classification model, and the classified input image can be output. The sub-processes illustrated in FIG. 3 are described in more detail below.FIG. 4 is a diagram illustrating a search operation for an image set 120 according to an embodiment of the present disclosure. As discussed above at 310, a plurality of images associated with a target text are captured via a semantic text-image matching model, the target text indicating a target scenario. Through this search process for the image set 120, the image set 120 may be roughly filtered to retrieve multiple images associated with the scenario of interest. Although there may be low noise in the multiple images retrieved, such coarse filtering has filtered out most irrelevant data and thus helped improve training of the classification model and reduce burden on subsequent processing.In some embodiments, obtaining the plurality of images associated with the target text may include using a text encoder to extract text features from the target text and using an image encoder to extract image features for each image in the image set 120. Then, the similarity between the extracted text features and the image features may be calculated, and a predetermined number of images whose calculated similarities of the image features are greater than a predetermined similarity threshold value is selected in the image set.In some embodiments, the semantic text-image matching model described above may include a pre-trained model that uses contrasting learning between text and image, and the text coder and the image coder may be pre-trained text coders and image coders included in such a pre-trained model that uses contrasting learning between text and image. The text-to-image semantic matching model can process text and image together, and by learning image and text together, the semantic associations between them are captured and then text-based image classification is performed. In some embodiments, the text-to-image semantic matching model may include a Contrastive Language-Image Pre-Training (CLIP) model.As shown in FIG. 4, at 410, the text-to-image semantic matching model is loaded. At 420, a text description of a scenario of interest is obtained and the text features extracted therefrom by a text encoder. In 430, a raw image is obtained and the image features are extracted therefrom by an image encoder. At 450, the similarity between the extracted text features and image features is calculated. At 450, a predetermined number of images are obtained based on the calculated similarity. For example, the first 100 images are obtained whose image features have calculated similarities that are greater than a predefined similarity threshold value. It should be understood that the predetermined number of examples described herein are for ease of understanding and explanation only and are not intended to be limiting.FIG. 5 illustrates the generation process of an image pattern set according to an embodiment of the present disclosure. According to the embodiments of the present disclosure, in order to automatically assign a corresponding category label to each of the searched images corresponding to the target scenario, images of the same category may be stored in the same order based on the image processing model, examples of image processing models including transformers (transformers), among others, and the image processing models may be pre-trained models. It should be understood that the examples of image processing models described herein are for the purpose of ease of understanding and ease of explanation only and are not limiting, and other different image processing models may also be used to process image pattern sets in the embodiments of the present disclosure.As shown in FIG. 5, at 510, the multiple images found are divided into different folders. At 520, a binary label is generated, meaning that each image in the multiple images found is assigned a positive or negative pattern category label, meaning that it is determined whether each image pattern is a positive or negative pattern. At 521, the patterns assigned a positive pattern category label are stored in the positive pattern folder, and at 522, patterns assigned a negative pattern category label are stored in the negative pattern folder.In 530, the generation of multiple category labels is performed, i.e. each of the searched images is assigned a corresponding category label from a plurality of category labels, which means determining the pattern category of each image pattern. Next, the three-piece license plate generation is used as an example of the multi-piece license plate generation, which is for the purpose of understanding and convenience of explanation only and is not limitative. In 531, the patterns assigned a label of pattern category A, for example, are stored in the order of pattern category A, in 532 patterns assigned a label of pattern category B are stored in the folder of pattern category B, and in 533 patterns assigned a label of pattern category C are stored in the folder of pattern category C, etc.In order for images subsequently input to a model having a format corresponding to the model, the images may be subjected to preprocessing. According to the embodiments of the present disclosure, the input size of the classification model may be adjusted by performing preprocessing for each image in the image set 120. The preprocessing used may include, but is not limited to, changing size and centered cropping. As shown in FIG. 5, the image to be edited is input at 540. At 550, the size of the image is adjusted and at 560 the image is cut centered. The preprocessed image may have, for example, a 224*224 or 336*336 pixel format, without being limited thereto.FIG. 6 illustrates the training process of a classification model according to an embodiment of the present disclosure. According to embodiments of the present disclosure, the classification model to be trained may include a linear acquisition model that uses contrasting learning between text and image. In some embodiments, an example of such a linear detection model that uses contrasting learning between text and image may include a pre-trained linear CLIP detection model and may include two fully connected layers to achieve a balance between speed and accuracy. It should be understood that the classification model to be trained may also include other different models to adapt to different actual situations and specific requirements, and that the model may include more or less fully linked layers, this disclosure not being limited in this respect.During the training process for the classification model, the image pattern set generated as described above may be used. Due to the use of a linear acquisition model that uses contrasting learning between text and image, the amount of data used during training and testing is relatively small compared to conventional methods, e.g. only a few dozen data patterns per category are required. The test pattern set may be larger to test the robustness and generalization capability of the classification model. It should be understood that more data is also possible, with a larger image pattern set being achievable by data augmentation. In other words, the classification model generation process 300 according to an embodiment of the present disclosure may be supervised learning or semi-supervised learning.In some embodiments, the image pattern set may include a training image pattern set having a first number of image patterns and a test image pattern set having a second number of image patterns, the first number being less than the second number. It should be understood that the number configuration of training patterns and test patterns described herein in the pattern set is exemplary only and that other number configurations may also be used.In the embodiments of the present disclosure, in the case of being divided into two categories in the image pattern set (i.e., positive pattern category and negative pattern category), a loss function for the difference between the predicted classification result and the corresponding category label for each image pattern in the image training pattern set may be used to train the classification model, wherein the loss function is based on a combination of sigmoid activation function and binary cross entropy loss calculation and focus loss (focal loss), and by minimizing the loss function, the parameter set of the classification model is adjusted. It should be noted that other suitable activation functions are also possible.As illustrated in FIG. 6, at 610, the pre-trained linear acquisition model is loaded between text and image using contrasting learning, and the system is switched to the training mode. At 620, a binary case is indicated as described above. At 640 and 650, BCEWithlogrtsloss a combination of binary cross entropic loss (binary cross entropic loss) and sigmoid activation function loss calculation and focus loss is used.Thus, due to binary classification, there is often an imbalance between the positive and negative pattern sets. According to the embodiments of the present disclosure, focus loss may be applied to BEWithLogites to focus the model more on difficult and misclassified examples by reducing the weight of simple examples as represented in Formula (1) below: where α controls the balance between simple and difficult examples. It is the weight that gives more importance to a few categories (positive categories) during the training process. The higher the value of α, the greater the weight of positive categories, which promotes the handling of class imbalance through greater attention in difficult positive patterns.Moreover, γ is another hyperparameter that controls the loss rate in well categorized examples. It introduces an adjustment factor that increases the loss over misclassified examples. Higher values of y increase the speed at which focus loss decreases in well-classified examples, emphasizing more difficult examples during training. It should be understood that other hyperparameters are also configurable to improve performance, e.g., number of epochs (Epoch), learning rate, stack size, etc.In this way, applying focus loss to BEWithLogites results in an improvement in BEWithLogites FL.In the embodiments of the present disclosure, when patterns in a pattern set are divided into multiple categories (e.g., but not limited to, the three categories pattern category A, B, and C), the classification model may be trained using a loss function for the difference between the predicted classification result and the corresponding category label for each pattern in the pattern set to train the classification model, the loss function based on a combination of sigmoid activation function and binary cross entropy loss calculation and focus loss, and by minimizing the loss function, the parameter set of the classification model is adjusted.As illustrated in FIG. 6, the situation of the multiple classification described above is displayed at 630, which is explained below in a non-limiting manner using the example of three classifications. In 650, 660, and 670, the cross entropy loss (CrossEntropyLoss), the smoothed cross entropy loss, and the focus loss are used. According to embodiments of the present disclosure, the cross entropy loss may be used first to calculate the cross entropy between predicted categories and true categories. Subsequently, smoothed cross entropy loss can be used, e.g., by replacing the "hard" labels 0 and 1 with slightly smoother values (such as 0.1 and 0.9).In some embodiments, a tensor may be created in which a smaller value is assigned to each category other than the target category. The target category may be assigned a higher probability and the remaining probability mass is distributed among the other classes. Next, the standard cross entropy loss can be calculated without label smoothing. Subsequently, the label smoothing effect is applied by element-by-element multiplication between the loss and the smoothed labels, so that the loss contribution of mispredicted categories based on the smoothed labels is reduced. Sum the losses and obtain a smoothed loss tensor for each pattern. Finally, the average of these losses is returned as the final smoothed loss value. In this way, the model can be made more robust by introducing the label smoothing into the prediction. Additionally or alternatively, the focus loss may also be applied to a plurality of classification situations as illustrated in the following formula (2): wherein α and γ may be formed the same as or as described above. The use of the cross entropy of the focus losses can further converge the losses to achieve higher accuracy, particularly in complex multi-classification tasks such as scenario classification of trucks with abnormal shapes. The final loss value can be reduced by an order of magnitude by using the cross entropy of the focus loss and the smoothing cross entropy.With a loss strategy as described above, the training of the linear capture may be performed at 680, and at 690, the linear capture model trained with the contrasting learning between text and image is stored. Therefore, by the training process of the classification model according to the embodiments of the present disclosure, the training time of the model is very fast, as e.g. only the last classification layer is trained for the target scenario. In particular, as compared to a few weeks in conventional methods, the training process in the embodiments of the present disclosure generally takes only a few minutes, which allows a rapid iteration of the classification model to find the best parameters.FIG. 7 illustrates the test process of a classification model according to an embodiment of the present disclosure. According to the embodiments of the present disclosure, a model accuracy for the classification model is obtained by testing the generated classification model using the test pattern set, and the obtained model accuracy is compared with a predefined accuracy threshold.As illustrated in FIG. 7, at 710, a linear sensing model trained between text and image for the target scenario is loaded and the system is switched to the test mode. At 720, the binary linear detection test is displayed and at 730, the multi-piece linear detection test is displayed. At 740 and 750, the test results are compared to the labels. In the case of binary linear detection, at 741, when the value indicated by the flag is 0, patterns having test values greater than or equal to zero are classified into the positive pattern category as positive patterns, and at 742, when the value indicated by the flag is 1, patterns having test values less than or equal to 0 are classified into the negative pattern category as negative patterns. Similarly, in the multi-component detection in steps 751, 752 and 753, the image patterns are classified into corresponding pattern categories such as pattern categories A, B and C based on the test values and the values indicated by the labels. At 760 and 770, the accuracy of the model is calculated for the test image set and the model accuracy is output. In some embodiments, the model accuracy score is determined based on, for example, the ratio of correctly classified image patterns to the overall patterns. At 780, the test result is output.According to the embodiments of the present disclosure, in response to the model accuracy obtained by the above-mentioned test process being below the predetermined accuracy threshold, the image pattern set may be updated by adding additional image patterns to the image pattern set, the additional image patterns being associated with a long-tail scenario, and the classification model may be re-trained using the updated image pattern set. In this way, if the training result is not ideal, by improving the pattern quality, a scenario with poor training effects can be retrained to obtain higher accuracy. It should be understood that the pattern quality can also be improved by removing noise, outliers, etc.FIG. 8 illustrates the prediction process of a classification model according to an embodiment of the present disclosure. According to the embodiments of the present disclosure, input images associated with a target scenario may be received, and a category to which each of the input images belongs is determined by a generated classification model, and the classified input images are output. A linear recognition model that uses contrasting learning between text and image trained for a target scenario by embodiments of the present disclosure is capable of providing accurate classification of input images via the target scenario. In some embodiments, the classified input images may be manually inspected and subjected to further labeling, then merged into a pattern set to form a closed loop and used for processing based on subsequent features.As illustrated in FIG. 8, at 810, a linear acquisition model under test trained between text and image for the target scenario using contrasting learning is loaded and the system is switched to the prediction mode. Predictions for binary linear detection are given in 820, and predictions for multiple linear detection are given in 830. At 840 and 850, inference features are obtained. In the case of binary linear detection, at 841, when the prediction value is larger than 0, the picture patterns are classified into the positive pattern category as positive patterns, and at 842, when the prediction value is smaller than 0, the picture patterns are classified into the negative pattern category as negative patterns.Similarly, in the multi-component detection in steps 851, 852 and 853, the image patterns are classified into respective pattern categories, e.g., pattern categories A, B and C, based on the prediction values. At 860, the predicted images and their categories are output for the target scenario.FIG. 9 shows a schematic diagram of an apparatus 900 for generating a classification model according to an exemplary embodiment of the present disclosure. The apparatus 900 may include multiple units or modules and is used to perform the corresponding steps in method 200 as discussed in FIG. 2. As shown in FIG. 9, the apparatus 900 includes: an image acquisition module configured to acquire a plurality of images related to a target text using a text-to-image semantic matching model, the target text indicating a target scenario; a pattern set generation module configured to generate an image pattern set including image patterns having category labels by determining the category to which each image of the plurality of images belongs; and a model generation module configured to generate a classification model for the target scenario based on the image pattern set, wherein the generation of the classification model is based on a pre-trained model using linear recognition.In some embodiments, the image acquisition module 910 is additionally configured to: extract text features from the target text using a text encoder; extract image features of each image in an image set using an image encoder; calculate the similarity between the extracted text features and image features; and select a predetermined number of images in the image set for which the calculated similarity of the image features is greater than a predetermined similarity threshold.In some embodiments, where, if the image set may include images captured by one or more vision sensors on the vehicle, the apparatus 900 may also include a pre-processing module configured to: adjust the input size of the classification model by pre-processing each image of the image pattern set, wherein the pre-processing includes the resize and a centered cropping.In some embodiments, the classification model may be based on contrasting learning and include two fully connected layers; the image pattern set may include a training image pattern set having a first number of image patterns and a test image pattern set having a second number of image patterns, the first number being less than the second number.In some embodiments, wherein if the number of category comprises 2 and the categories may comprise positive pattern categories and negative pattern categories, the model generation module 930 may additionally be configured to: train the classification model using a loss function of the differences between the predicted classification results and the corresponding category label for each image pattern in the image training pattern set, the loss function based on a combination of an activation function with a binary cross entropy loss calculation and a focus loss; and adjust a parameter set of the classification model by minimizing the loss function.In some embodiments, the number of categories may include three or more than three, and the model generation module 930 may be further configured to: train the classification model using a loss function of the differences between the predicted classification results and the corresponding category label for each image pattern in the image training pattern set, the loss function based on a cross entropy loss, a smoothed cross entropy loss, and a focus loss; and adjust a parameter set of the classification model by minimizing the loss function.In some embodiments, the apparatus 900 may further include a test module configured to: obtain a model accuracy score for the classification model by testing the generated classification model using the test pattern set; and compare the obtained model accuracy score with a predetermined accuracy threshold.In some embodiments, the apparatus 900 may further include a retraining module configured to respond to a model accuracy less than the predetermined accuracy threshold, update a sample set by adding additional samples associated with a long-tail scenario; and perform retraining the classification model using the updated sample set.In some embodiments, the target scenario may include tunnel inputs and the categories may include inputs and non-inputs; or the target scenario may include tunnel environments and the categories may include inputs, outputs, interiors, and others.FIG. 10 shows a schematic block diagram of an example apparatus 1000 that may be used to implement an embodiment of the present disclosure. As shown in the figure, the apparatus 1000 comprises a central processing unit (CPU) 1001 which can perform various suitable actions and processes in accordance with computer program instructions stored in a read only memory (ROM) 1002 or loaded from a storage unit 1008 into a random access memory (RAM) 1003. The RAM 1003 may further store various programs and data required for the operation of the apparatus 1000. The CPU 1001, ROM 1002 and RAM 1003 are connected to each other via a bus 1004. An input / output interface (I / O) 1005 is also connected to bus 1004.Multiple components in device 1000 are connected to I / O interface 1005, including: an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of display devices, speakers, etc.; a memory 1008, such as a hard disk, optical memory, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices via computer networks, such as the Internet and / or various telecommunications networks.The various methods and processes described above, such as the method 200 and the process 300 and its sub-processes, may be executed by the processing unit 901. In some embodiments, the method 200 and process 300 as well as its sub-processes may be implemented, for example, as a computer software program that is physically contained in a machine readable medium such as the storage unit 1008. In some embodiments, the computer program may be fully or partially loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the CPU 1001, one or more of the above-described actions of the method 200 and the process 300 and its sub-processes may be performed. According to embodiments of the present disclosure, a vehicle is provided that may include an apparatus 900 as described above for carrying out various aspects of the present disclosure.The present disclosure may be a method, apparatus, electronic device, vehicle, computer readable storage medium, and / or computer program product. The computer program product may include a computer readable storage medium having thereon computer readable program instructions for carrying out the various aspects of the present disclosure.A computer readable storage medium may be a tangible device that can receive and store instructions for use by an apparatus for instruction execution. The computer readable storage medium may be, for example, but not limited to, an electronic storage medium, a magnetic storage medium, an optical storage medium, an electromagnetic storage medium, a semiconductor storage medium, or a suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer readable storage media include: portable computer diskettes, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or Flash memories), static random access memories (SRAM), portable compact disk read-only memories (CD-ROM), digital versatile disks (DVD), memory sticks, floppy disks, mechanically encoded devices such as punch-cards or raised structures in grooves on which instructions are stored, and any suitable combination of the foregoing. The computer readable storage media used herein are not to be interpreted as transient signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through an optical fiber cable), or electrical signals transmitted through wires.The computer readable program instructions described herein may be downloaded from a computer readable storage medium to various computer / processing devices or to an external computer or external storage device via a network such as the Internet, a local area network or a wide area network, and / or a wireless network. A network adapter card or network interface in each computer / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage on a computer readable storage medium in the respective computer / processing device.The computer program instructions for carrying out operations within the scope of the present disclosure may be assembler instructions, ISA instructions (instruction set architecture), machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, the programming languages including object oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code may execute entirely on a computer, partly on a computer, as a stand-alone software package, partly on a computer and partly on a remote computer or entirely on a remote computer or server. In the latter case, the remote computer may be connected to a computer by any type of network (including a local area network (LAN) or a wide area network (WAN)), or may be connected to an external computer (for example, by an Internet Service Provider via the Internet). In some embodiments, the electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be adapted using state information of computer readable program instructions, wherein the electronic circuit may execute the computer readable program instructions to implement various aspects of the present disclosure.Aspects of the present disclosure have been described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, may be implemented by computer readable program instructions.These computer readable program instructions may be provided to a processing unit of a general purpose computer, special purpose computer, or other programmable data processing apparatus, thereby generating a machine such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, generate means for implementing the functions / acts indicated in the flowchart and / or block diagram. These computer readable program instructions may also be stored in a computer readable storage medium that causes a computer, a programmable data processing device, and / or other device to function in a particular manner, such that the computer readable medium having the instructions stored thereon comprises an article of manufacture including instructions that implement various aspects of the functions / acts included in the flowchart and / or block diagram block or blocks.Computer readable program instructions may also be loaded into a computer, other programmable data processing device, or other device that performs a series of operations on a computer, other programmable data processing device, or other device to produce a computer implemented process such that the instructions executed on the computer, other programmable data processing device, or other device implement the functions / acts specified in one or more K blocks in the flowchart and / or block diagram.The flow diagrams and block diagrams in the figures illustrate the functions and operations that may be implemented in accordance with the method of the embodiments. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions that includes one or more components for implementing the specified logical function(s). In some alternative implementations, the functions indicated in the appended blocks may be performed in a different order than in the figures. For example, two successive blocks may be executed substantially in parallel, and may sometimes be executed in the reverse order depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a special purpose hardware-based system that performs the specified function or action, or may be implemented by a combination of special purpose hardware and computer instructions.The embodiments of the present disclosure have been described above, and the above description is illustrative, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to best explain the principles, practical applications, or technical improvements in technology on the market for the embodiments, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
A method of generating a classification model, comprising: acquiring a plurality of images associated with a target text using a text-to-image semantic matching model, the target text indicating a target scenario; generating an image pattern set comprising image patterns having category labels by determining the category to which each image of the plurality of images belongs; and generating a classification model for the target scenario based on the image pattern set, wherein generating the classification model is based on a pre-trained model using linear recognition.The method of claim 1, wherein acquiring multiple images associated with a target text comprises: extracting text features from the target text using a text encoder; extracting image features of each image in an image set using an image encoder; calculating similarity between the extracted text features and image features; and selecting a predetermined number of images in the image set for which the calculated similarity of the image features is greater than a predetermined similarity threshold.The method of claim 2, wherein the image set comprises images captured by one or more vision sensors on the vehicle, the method further comprising: performing preprocessing for each image in the image set to adjust the input size of the classification model, wherein the preprocessing comprises changing the size and centered cropping.The method of claim 1, wherein: the classification model is based on comparative learning and comprises two fully connected layers; and the image pattern set comprises a training image pattern set having a first number of the image patterns and a test image pattern set having a second number of the image patterns, the first number being less than the second number.The method of claim 4, wherein the number of categories comprises 2 and the categories comprise positive pattern categories and negative pattern categories, and wherein generating the classification model comprises: training the classification model using a loss function of the differences between the predicted classification results and the corresponding category label for each image pattern in the image training pattern set, wherein the loss function is based on a combination of an activation function with a binary cross entropy loss calculation and a focus loss; and adjusting a parameter set of the classification model by minimizing the loss function.The method of claim 4, wherein the number of categories comprises three or more than three, and wherein generating the classification model comprises: training the classification model using a loss function of the differences between the predicted classification results and the corresponding category label for each image pattern in the image training pattern set, wherein the loss function is based on a cross entropy loss, a smoothed cross entropy loss, and a focus loss; and adjusting a parameter set of the classification model by minimizing the loss function.The method of claim 5 or 6, further comprising: obtaining a model accuracy score for the classification model by testing the generated classification model using the test image pattern set; and comparing the obtained model accuracy score to a predefined accuracy threshold.The method of claim 7, further comprising: in response to the model accuracy score being less than the predetermined accuracy threshold, updating the image pattern set by adding additional image patterns to the image pattern set, the added image patterns associated with a long-tail scenario; and retraining the classification model using the updated image pattern set.The method of claim 1, further comprising: receiving input images associated with the target scenario; determining the category to which each input image among the input images belongs using the generated classification model; and outputting the classified input images.The method of claim 1, wherein the target scenario comprises a tunnel input and the categories comprise an input and a non-input; or the target scenario comprises a tunnel environment and the categories comprise an input, an output, an interior, and others.An apparatus for generating a classification model, comprising: an image acquisition module configured to acquire a plurality of images associated with a target text using a semantic text-image matching model, the target text indicating a target scenario; a pattern set generation module configured to generate an image pattern set comprising image patterns having category labels by determining the category to which each image of the plurality of images belongs; and a model generation module configured to generate a classification model for the target scenario based on the image pattern set, wherein the generating of the classification model is based on a pre-trained model using linear recognition.An electronic device, comprising: at least one processor; and a memory coupled to the at least one processor and having instructions stored thereon that, when executed by the at least one processor, cause the device to perform a method according to any one of claims 1 to 11.A vehicle comprising the electronic device of claim 12.A computer readable storage medium having computer executable instructions stored thereon, the computer executable instructions being executed by a processor to implement a method according to any one of claims 1 to 12.
Citation Information
Cited By
Text-to-image generation model control method and device and storage medium
CN121328775A
Method, device and storage medium for controlling a text-to-image generation model
CN121328775B