Picture data preprocessing method and device, equipment, medium and product
By preprocessing the image data into a feature vector and storing it in the vector database, the problems of repeated operations and high complexity in the preprocessing process of image data in the prior art are solved, and the effect of reducing model training costs and resource consumption is achieved.
Patent Information
- Application Number
- CN202510686648.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The prior art has repeated operations and high time and space complexity in the preprocessing of picture data, resulting in increased model training costs and resource consumption.
By directly storing image data in the vector database in the form of feature vectors, automatic filtering, data equalization and feature dimension optimization are achieved, repeated operation processes are avoided, and the time and spatial complexity of model training are reduced.
It reduces the hardware resources required during model training and parameter adjustment, reduces the cost of model training, and improves resource utilization efficiency.
Smart Images

Figure CN120219887A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image processing in artificial intelligence, and particularly to a method, device, equipment, medium and product for preprocessing picture data. Background Art
[0002] When detecting and recognizing target objects in pictures, the quality and data volume of the image itself will directly affect the design of the subsequent model architecture, the selection of loss functions, the implementation of optimization algorithms, and the accuracy of recognition effects, etc. Therefore, before performing detection or recognition operations, appropriate preprocessing operations need to be performed on the pictures. Picture preprocessing is mainly to eliminate some noise information in the pictures, enhance the readability of effective information, optimize data information to a greater extent, and provide good conditions for the detection of target objects in subsequent pictures. Existing data preprocessing technologies generally check for abnormalities in the data, perform data denoising according to business requirements, remove useless or illegal information, and then hand over all the graphic data (text data is the label of the picture data) to the model for training. After new business requirements emerge, developers need to perform data preprocessing again according to the new business logic, and then use the graphic data as input for model training. Inside the model, vectorization is performed to convert the graphic data into feature vector data. Such a process increases the time cost of data preprocessing. And a large amount of graphic data contains data irrelevant to business requirements, and these irrelevant data will further increase the training cost and space complexity of the model. When adjusting the algorithm of the model or wanting to adjust algorithm parameters, this preprocessing process will be repeated. Summary of the Invention
[0003] The purpose of this application is to provide a method, device, equipment, medium and product for preprocessing picture data, which stores picture data directly in the form of feature vectors in a vector database. When used for model training, it can not only avoid repeated operation processes, but also reduce the time complexity and space complexity of model training, and reduce the hardware resources required during model training and parameter tuning under the same conditions.
[0004] To achieve the above purpose, the following solutions are provided in this application.
[0005] In the first aspect, this application provides a method for preprocessing picture data, including: Obtain the corresponding image dataset according to the task objectives of the object detection model; the task objectives of the object detection model include human object detection, animal object detection, vehicle object detection, lesion area detection, obstacle detection, product defect detection, abnormal item detection, and pest and disease area detection; the corresponding image datasets include human image dataset, animal image dataset, vehicle image dataset, lesion image dataset, obstacle image dataset, product image dataset, item image dataset, and crop image dataset; Classify and screen various image data in the image dataset according to the target types to be detected to obtain the target type dataset; Perform data balancing processing on the image data of each target type in the target type dataset to obtain the balanced dataset; Perform feature dimension optimization processing on each image data in the balanced dataset to obtain the optimized feature vectors; Store the optimized feature vectors in the vector database for use in training the object detection model.
[0006] Optionally, the classifying and screening various image data in the image dataset according to the target types to be detected to obtain the target type dataset specifically includes: Calculate the similarity between each image data in the image dataset and each target type through the CLIP model, and retain the image data with a similarity greater than the similarity threshold to form the target type dataset.
[0007] Optionally, the performing data balancing processing on the image data of each target type in the target type dataset to obtain the balanced dataset specifically includes: Judge whether the ratio of the image data between any two target types in the image data of each target type in the target type dataset is balanced; If the ratio is balanced, compress the image data of each target type to the same classification quantity Nmax; where Nmax is the number of images contained in the target type with the largest proportion after data compression; If the ratio is unbalanced, perform data augmentation on the image data of the target type with a smaller proportion, perform data compression on the image data of the target type with a larger proportion, and enhance or compress the image data of each target type to the same classification quantity Nmax.
[0008] Optionally, the performing feature dimension optimization processing on each image data in the balanced dataset to obtain the optimized feature vectors specifically includes: Perform block processing on each image data in the balanced dataset, and divide each image data into multiple image blocks of M rows × N columns; both M and N are integers greater than 1; Perform block-level classification on each image block, mark the image blocks containing the target object as 1, and mark the image blocks not containing the target object as 0 to obtain labeled image blocks; Based on the labels carried by each image block, screen and reorganize the multiple image blocks obtained by splitting each image data to obtain an optimized feature vector.
[0009] Optionally, the step of screening and reorganizing the multiple image blocks obtained by splitting each image data based on the labels carried by each image block to obtain an optimized feature vector specifically includes: Traverse the multiple M×N image blocks obtained by splitting each image data. Taking each row as a unit, retain the image blocks with label 1 and discard the image blocks with label 0; Reorganize all the image blocks with label 1 into a feature matrix with a dimension of 1×X. After the reorganization of each row of image blocks is successful, add the sequence start position marker [b], and add the sequence end position marker [e] after the reorganization of the last row of image blocks to obtain an optimized feature vector; where X is the number of image blocks with label 1 in each image data.
[0010] Optionally, the step of storing the optimized feature vector in a vector database for use in training a target detection model specifically includes: Use the optimized feature vector as a feature data set and store it in the vector database; When training the target detection model, divide the feature data set into a training set, a validation set, and a test set for training, validating, and testing the target detection model.
[0011] In a second aspect, the present application provides an image data preprocessing device, including: An image data set acquisition module, configured to obtain a corresponding image data set according to the task objective of the target detection model; the task objectives of the target detection model include human object detection, animal object detection, vehicle object detection, lesion area detection, obstacle detection, product defect detection, abnormal item detection, and pest and disease area detection; the corresponding image data sets include human image data sets, animal image data sets, vehicle image data sets, lesion image data sets, obstacle image data sets, product image data sets, item image data sets, and crop image data sets; An image classification and screening module, configured to classify and screen various image data in the image data set according to the target types to be detected to obtain a target type data set; A data balance processing module, configured to perform data balance processing on the image data of each target type in the target type data set to obtain a balanced data set; A feature dimension optimization module, configured to perform feature dimension optimization processing on each image data in the balanced data set to obtain an optimized feature vector; A feature vector storage module for storing the optimized feature vectors in a vector database for use in training a target detection model.
[0012] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the picture data preprocessing method.
[0013] In a fourth aspect, the present application provides a computer-readable storage medium with a computer program stored thereon, and when the computer program is executed by a processor, it implements the picture data preprocessing method.
[0014] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the picture data preprocessing method.
[0015] According to the specific embodiments provided by the present application, the following technical effects are disclosed.
[0016] A picture data preprocessing method, device, equipment, medium and product provided by the present application, on the one hand, automatically screens the target type data set according to the task objective of the target detection model and the target types to be detected and balances the type proportions, and on the other hand, performs feature dimension optimization processing on the picture data to achieve feature dimensionality reduction. The picture data preprocessed by the method of the present application is directly stored in the vector database in the form of feature vectors. In this way, when used for training a target detection model, it not only avoids repeated operation processes, but also reduces the time complexity and space complexity of the model, reduces the training time of the model under the same conditions, and reduces the hardware resources required during the training and parameter tuning processes, thereby effectively reducing the model training cost. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 It is a schematic flowchart of a picture data preprocessing method of the present application; Figure 2 It is a schematic diagram of the original picture of a tiger in an embodiment of the present application; Figure 3 It is a schematic diagram of block processing of the original picture in an embodiment of the present application; Figure 4Schematic diagram for block - level classification of picture blocks in the embodiments of this application; Figure 5 Schematic diagram for screening and reorganizing picture blocks in the embodiments of this application. Detailed implementation manners
[0019] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0020] This application proposes a method, device, equipment, medium and product for pre - processing picture data. The aim is to directly store picture data in the vector database in the form of feature vectors through an improved data pre - processing strategy. When used for model training, it can not only avoid repeated operation processes, but also reduce the time complexity and space complexity of model training, and reduce the hardware resources required during the model training and parameter tuning processes under the same conditions.
[0021] To make the above - mentioned objects, features and advantages of this application more obvious and understandable, the following further describes this application in detail in conjunction with the accompanying drawings and specific implementation manners.
[0022] In an exemplary embodiment, as Figure 1 shown, a method for pre - processing picture data is provided, including the following steps 1 to 5.
[0023] Step 1: Obtain the corresponding picture data set according to the task objective of the object detection model.
[0024] Object detection is a core task in computer vision, aiming to identify and locate target objects in an image (or picture), and generate bounding boxes and class labels for each detected target object. Object detection technology is widely used in multiple fields. For example, in the field of security monitoring, it is used to analyze surveillance videos in real - time to detect abnormal behaviors or targets, such as human objects or detection of abnormal items. For example, in the field of intelligent transportation, in autonomous driving and traffic management, it detects and tracks targets such as vehicles and pedestrians to achieve traffic congestion prediction and intelligent driving. For example, in the field of intelligent retail, it is used for shelf commodity detection and customer behavior analysis to improve retail efficiency and customer experience. Also, for example, in the field of medical image analysis, it assists doctors in lesion detection, cell recognition, etc., to improve the accuracy and efficiency of diagnosis.
[0025] Therefore, the task objectives of the object detection model described in this application can be human object detection, animal object detection, vehicle object detection, lesion area detection, obstacle detection, product defect detection, abnormal item detection, and pest and disease area detection, etc. The corresponding acquired image datasets should be human image datasets, animal image datasets, vehicle image datasets, lesion image datasets, obstacle image datasets, product image datasets, item image datasets, or crop image datasets.
[0026] Step 2: Classify and filter various image data in the image dataset according to the types of objects to be detected to obtain an object type dataset.
[0027] The types of objects to be detected correspond to the task objectives of the object detection model. For the convenience of subsequent description, the object detection model can be simply referred to as the model, and the word "type" can also be referred to as "category". For example, for lesion area detection, the types of objects to be detected can be lesion areas such as tumors and nodules, and normal tissue areas. For example, for obstacle detection, the types of objects to be detected can be pedestrians, other vehicles, and various objects on the road in front of the vehicle. For example, for product defect detection, the types of objects to be detected can be product surface defect types such as cracks, scratches, bubbles, and spots. For example, for abnormal item detection, the types of objects to be detected can be suspicious objects such as left-behind packages and dangerous items. For pest and disease area detection, the types of objects to be detected can be pest and disease types and crop areas affected by pests and diseases, etc.
[0028] The initially acquired image dataset contains various image data, and the image data can also be simply referred to as images. The purpose of Step 2 is to extract images that meet the model task objectives through an automated method to form an object type dataset.
[0029] The CLIP (Contrastive Language-Image Pretraining) model is a multi-modal machine learning model proposed by OpenAI in 2021. The core idea of the CLIP model is to map images and texts to the same semantic space through large-scale image-text pair training to achieve cross-modal understanding. The training data of the CLIP model is 400 million pairs of publicly available (image, text) data on the Internet. The CLIP model has the following capabilities: zero-shot classification, which can directly classify images according to text descriptions without fine-tuning for specific tasks. The characteristics of the CLIP model include: 1) No training data is required: it directly supports zero-shot classification; 2) Dynamically adjust categories: can modify candidate category texts at any time without retraining; 3) Cross-domain generalization: applicable to open-domain scenarios, having been exposed to a large amount of open-domain data during training and being able to understand a wide range of topics, such as objects, scenes, and abstract concepts.
[0030] The inputs of the CLIP model include image input and text input. Among them, the image input refers to a single picture, such as an RGB image. In its preprocessing process, the image needs to be resized to a fixed resolution (such as 224×224) and the pixel values are normalized. The Vision Transformer (ViT) or ResNet encoder is used to extract image features. The text input refers to a category or sentence described in natural language, such as "a picture of a dog". Its preprocessing process needs to convert the text into a sequence of tokens through a tokenizer. The Transformer encoder is used to extract text features. The output of the CLIP model is the similarity score between the input image and text. The CLIP model calculates the cosine similarity (range: real number, no fixed upper and lower bounds) between the image features and text features, and after normalization by Softmax, obtains the matching degree, that is, the similarity, between each text description and the image.
[0031] Therefore, when calculating the similarity between each picture data in the picture dataset and each target category in this application, only any candidate category text (such as target categories like cracks, scratches, bubbles, spots, etc.) and any picture data in the picture dataset need to be input, and the CLIP model can calculate the similarity between each picture data and each target category text. Further, the picture data with a similarity greater than the similarity threshold is retained to form the target category dataset. The range of the similarity threshold is between 0 and 1. The closer it is to 1, the greater the possibility of being the target category. The threshold can be set according to the actual situation. For example, the similarity threshold is set to 0.75.
[0032] The following uses a specific example for illustration. For example, in an exemplary embodiment, the task objective of the object detection model is to perform animal object detection, and the corresponding image dataset obtained includes various animal images such as lions, tigers, elephants, monkeys, cheetahs, zebras, penguins, bears, etc. And the target objects of the object detection model that the user wants to establish are lions, tigers, and cheetahs, then the target categories are determined to be lions, tigers, and cheetahs. Then the purpose of step 2 is to classify each image in the image dataset into one of the categories of lions, tigers, and cheetahs. Then the specific classification and screening process includes: 1) Text encoding: Convert each target category into a text description (such as "picture of a lion", "picture of a tiger", and "picture of a cheetah"), and encode it into a corresponding text feature vector; 2) Image encoding: Encode the image to be classified into an image feature vector; 3) Similarity calculation: Calculate the cosine similarity between the image feature vector and the three text feature vectors respectively; 4) Probability normalization: Obtain the similarity probability distribution through Softmax, and select the category with the highest probability as the image label. Calculate the similarity between each image and the given 3 types of labels (lions, tigers, cheetahs), and retain the image samples with similarity > 0.75, and mark the rest as "Other". Finally, in the target category dataset, there are pictures of the three animals, lions, tigers, and cheetahs, and each picture will have label information of at least one of the three animals.
[0033] Step 3: Perform data balancing processing on the image data of each target category in the target category dataset to obtain a balanced dataset.
[0034] In the target category dataset obtained in step 2, there are images of multiple target categories. In model training, the more balanced the proportion of each category in the training data, the better the training effect of the model. If the proportions of each category are very different, then the trained model may be "biased": for the images with a larger overall proportion, the training effect of the model will be much better than that of the images with a smaller proportion. So in this step, it will be judged whether the data proportions of the images of the three categories, lions, tigers, and cheetahs, in the above specific example are balanced. For example, if the ratio between any two of the data of each category in the images containing the above 3 target categories is 1:1, that is, the ratio of the number of images of the 3 target categories is 1:1:1, it means that the data is very balanced. The specific steps of step 3 include the following steps 3.1 to 3.3.
[0035] Step 3.1: Judge whether the ratio of the image data between any two target categories in the image data of each target category in the target category dataset is balanced.
[0036] The most ideal state of data ratio balance is that the number of pictures for each target category is the same. Therefore, for the target category dataset after classification and screening, first determine whether the picture data ratio between any two target categories in the picture data of each target category is balanced, that is, determine whether the quantity ratio between any two pieces of picture data of each category is 1:1.
[0037] Step 3.2: If the ratio is balanced, compress the picture data of each target category to the same classification quantity Nmax; where Nmax is the number of pictures included in the target category with the largest proportion after data compression.
[0038] If the data ratio is balanced, that is, the picture data ratio between any two target categories is 1:1, then directly compress the picture data of each target category. During the compression process, based on the number of pictures included in the target category with the largest proportion after compression (referred to as the classification quantity) Nmax, compress the number of pictures of the remaining target categories to Nmax as well.
[0039] The method of data compression is to calculate the cosine similarity based on the feature vectors in the pictures, retain the pictures with higher similarity, and give priority to retaining the pictures with higher model thresholds, thereby reducing data duplication and promoting the improvement of data quality. The calculation formula of cosine similarity is as follows: (1); Among them, is the image feature vector, which is extracted by the image encoder of the CLIP model (such as ViT or ResNet). The CLIP model converts the input image into a high-dimensional vector of a fixed dimension. For example, the output of CLIP-ViT is a 512-dimensional or 768-dimensional vector, and this vector encodes the global semantic information of the image, such as object category, scene, texture, etc. is the text feature vector, which is extracted by the text encoder of the CLIP model (such as Transformer). The input text (such as "tiger") is mapped by the text encoder into a vector of the same dimension as the image feature (such as 512 dimensions). The text feature vector encodes the semantic information of the label. For example, "tiger" will be associated with concepts such as "animal", "mane", "forehead stripes", "grassland", etc.
[0040] represents the square root of the sum of the squares of each element of the vector, that is: (2); Among them, represents the vector of the -th dimension element, is the vector dimension. For example, for a 512-dimensional feature vector, in the formula= 512. The function of formula (2) is to eliminate the influence of vector length on similarity in non - normalized scenarios and only retain the direction information. In the CLIP model, L2 normalization is performed on image and text features, enabling the cosine similarity to be directly calculated through the dot product. That is, in actual calculation, the denominator in formula (1) can be omitted, and there is no need to calculate formula (2) anymore.
[0041] That is to say, the denominator in formula (1) is the product of the L2 norms of the two vectors, aiming to normalize the dot product to the interval [-1, 1]. The CLIP model used in this application performs L2 normalization on image and text features during training, that is, it forces . Therefore, in the inference stage, the encoded feature vectors already automatically meet the normalization condition, and the value of the denominator is: . At this time, the cosine similarity is simplified to: (3); That is, there is no need to explicitly divide by the denominator, and the dot product can be directly used to obtain the cosine similarity. The normalized dot product is equivalent to the cosine similarity, with higher calculation efficiency and consistent with the model training objective.
[0042] The following takes the classification of "tiger" pictures as an example for introduction. Image encoding: Use the CLIP model to encode the picture into a 512 - dimensional vector . Text encoding: Encode the label "tiger" into a vector of the same dimension . Similarity calculation: . Then when there are multiple pictures with similarities to the label "tiger" between 0.84 and 0.85 (the interval range [0.84, 0.85] here is a small interval), data compression will retain the pictures with an image similarity calculation of 0.85.
[0043] Of course, regardless of the value of the calculated similarity, it is similar to the data compression method within the interval [0.84, 0.85]. Moreover, when dividing small intervals within the similarity interval where all the pictures of the target category (tiger) are located, the size of the small interval is determined according to the actual situation. Suppose for all the pictures of the target category "tiger", the similarity of all the pictures is within the interval [0.86, 0.96], then the span of the small interval during data compression is set to 0.005. That is, for the pictures within the interval [0.86, 0.865] of one of the small intervals, the picture with a similarity of 0.865 is retained. Suppose for all the pictures of the target category "tiger", the similarity of all the pictures is within the interval [0.87, 0.90], that is, the range of the similarity interval is relatively small, then the span of the small interval during data compression is set to 0.001. That is, for the pictures within the interval [0.87, 0.871] of one of the small intervals, the picture with a similarity of 0.871 is retained. And so on, the compression of the number of pictures can be achieved.
[0044] Step 3.3: If the proportion is unbalanced, perform data augmentation on the picture data of the target category with a smaller proportion, and perform data compression on the picture data of the target category with a larger proportion, and enhance or compress the picture data of each target category to the same classification quantity Nmax.
[0045] If the data proportion is unbalanced, then perform data compression on the picture data with a larger proportion. During the compression process, based on the largest classification quantity Nmax after compression, the quantities of the other classifications with larger proportions are also compressed to Nmax. When calculating each type of picture using cosine similarity, the obtained interval ranges may be different, that is, the number of retained pictures will be different. Nmax can only be obtained after all the picture data is compressed, and then take the maximum value among them. The purpose of doing this is to retain the diversity of scarce data.
[0046] Conversely, perform data augmentation on the picture data with a smaller proportion, and the number of pictures after augmentation is also Nmax. Data augmentation can adopt processing methods such as image rotation, cropping, translation, and flipping.
[0047] Step 4: Perform feature dimension optimization processing on each picture data in the balanced dataset to obtain optimized feature vectors.
[0048] The said Step 4 specifically includes the following Steps 4.1 to 4.3.
[0049] Step 4.1: Perform block processing on each picture data in the balanced dataset, and divide each picture data into multiple picture blocks of M rows × N columns.
[0050] The picture is cut into blocks in order to subsequently identify whether each picture block contains the target object to be recognized. Further, each picture block can be flattened into a vector, and after linear projection, its embedding representation is obtained, and positional encoding is added to retain spatial information. Among them, flattening each picture block into a vector and obtaining its embedding representation is to facilitate the model's calculation and recognition. The picture information seen by the eyes needs to be processed into a special form of digital expression that can be processed by a computer. Adding positional encoding is to retain the spatial information between the picture blocks obtained by splitting the original picture. After the original picture is segmented, the overall dimension is M×N; both M and N are integers greater than 1. The original dimension of the picture is height multiplied by width. For the sake of convenient expression, h represents the height and w represents the width, that is: h×w. Here, the pixel size of the segmented picture blocks can be adjusted according to the actual situation, for example, taking 16×16 pixels.
[0051] The following uses a specific example to illustrate. For Figure 2 the picture of the "tiger" category shown, use ViT to segment the original picture into picture blocks of a fixed size (such as 16×16 pixels). As Figure 3 shown, the picture is segmented into 4 rows × 3 columns, a total of 12 picture blocks. Figure 2 The original overall dimension of the tiger picture in Figure 3 is (4×16)×(3×16). After segmentation,
[0052] the overall dimension of Figure 3 is M×N = 4×3.
[0052] Further, each picture block can be flattened into a vector. Assume that the original picture size is 256×256 pixels and it is an RGB three-channel image, then its actual shape is 3×256×256. The block size is set to 16×16 pixels, that is, the actual shape of each picture block is 3×16×16. The flattening operation of the picture block means that each 16×16-sized picture block is regarded as an independent small image, and the 16×16 pixel values of the R / G / B three channels of the picture block are arranged in order into a one-dimensional vector, so the vector length of each picture block is 3×16×16 = 768.
[0053] Assume that for a 16×16 three-channel picture block, the implementation code example is as follows: block = torch.randn(3, 16, 16) # Shape [3, 16, 16] flattened_block = block.flatten( ) # Shape
[768] Among them, torch.randn( ) is a function in PyTorch used to generate random tensors. block represents an image patch. block.flatten( ) refers to the flattening operation of the image patch. flattened_block refers to the flattened vector.
[0054] Furthermore, the flattened 768-dimensional vector is mapped to the unified dimension required by ViT (such as 512 dimensions) in the following way. Use a learnable fully connected layer (linear layer) with a weight matrix shape of 768×512. Multiply each flattened vector (768 dimensions) by this matrix, and the output is a 512-dimensional embedded vector. This process is called "linear projection", which essentially compresses and abstracts the original pixel information. The implementation code example is as follows: linear_layer = torch.nn.Linear(768, 512) # Define the linear layer embedded_vector = linear_layer(flattened_block) # Output shape
[512] Among them, torch.nn.Linear( ) is a very important module in PyTorch used to implement a fully connected layer (also known as a linear layer). The fully connected layer is a common layer type in neural networks used to convert input data into output data, usually for processing one-dimensional data. linear_layer refers to the generated fully connected layer. embedded_vector is the obtained embedded vector.
[0055] Furthermore, position encoding is added to the embedded vector to retain spatial information. The implementation code of the relevant method has been built into the Transformer library, and the code for calling the relevant function can achieve it.
[0056] Step 4.2: Perform block-level classification on each image patch, mark the image patch containing the target object as 1, and mark the image patch not containing the target object as 0 to obtain the labeled image patches.
[0057] This application performs block-level classification on each image patch by fine-tuning the ViT structure. Specifically, it is achieved by adding a classification head (fully connected layer + Softmax) to the output embedding of each block after the Transformer encoder of ViT.
[0058] The main modules of ViT include: image patching and embedding (Patch Embedding), positional encoding (Positional Encoding), Transformer encoders (multiple layers), and a classification head (usually a fully connected layer). In this application, after the Transformer encoder of ViT, a classification head (fully connected layer + Softmax) is added to the output embedding of each patch. The role of the fully connected layer is to map the high-dimensional embedding of each patch output by the Transformer to the category space (such as a two-dimensional space, corresponding to 0 and 1). Softmax then converts the linear output into a probability distribution, representing the probability that each patch belongs to each category. The input of the fully connected layer is the embedding vector of each patch (embedding_dim, such as 512 dimensions), and the output is the dimension of the number of categories (such as 0 and 1, two dimensions). The input of Softmax is this two-dimensional vector, and the output is a normalized probability distribution.
[0059] Input each image patch into the improved ViT structure. After block-level classification, mark the image patches containing the target object as "1", which are called target patches; mark the image patches not containing the target object as "0". "0" and "1" are the category labels of the image patches. For example Figure 4 as shown in, mark the image patches containing the target object "tiger" as "1", and mark the image patches not containing the target object "tiger" as "0".
[0060] Step 4.3: Based on the labels carried by each image patch, screen and reorganize the multiple image patches obtained by splitting each image data to obtain an optimized feature vector.
[0061] Specifically, traverse the M rows × N columns of multiple image patches obtained by splitting each image data. Taking each row as a unit, retain the image patches with the label "1" (i.e., target patches) and discard the image patches with the label "0". Further, reorganize all the image patches with the label "1" into a one-dimensional feature matrix of 1×X dimensions, and add a sequence start position marker [b] after each row of image patches is successfully reorganized, and add a sequence end position marker [e] after the reorganization of the last row of image patches, so as to obtain an optimized feature vector, that is, obtain a sequence composed of target patches and markers [b] and [e]. Where X is the number of image patches with the label "1" in each image data. [b] represents beginning, which is used to indicate the start position of the sequence. [e] represents ending, which is used to indicate the end position of the sequence.
[0062] As Figure 4 and Figure 5As shown, after the image of the tiger is block-recognized, the image blocks in the first and second rows both contain part of the information of the tiger, so they are marked as 1. Then, when reorganizing, these two rows are retained. When reorganizing, add [b] after the end of the image block in the first row, and then add the image block in the second row. Add [b] after the end of the image block in the second row. The first image block in the third row does not contain any part of the information of the tiger, so it is marked as 0. Discard this image block with a label of 0, and add the remaining two image blocks marked as 1 and a [b] in the third row to the previous [b]. The first image block in the fourth row also does not contain any part of the information of the tiger, so it is marked as 0. Discard this image block with a label of 0, and add the remaining two image blocks marked as 1 and an [e] in the fourth row to the previous [b], because this is the last image block of the image and the sequence ends.
[0063] Step 5: Store the optimized feature vectors in the vector database for use in training the target detection model.
[0064] After the image data undergoes feature dimension optimization in Step 4, it is now represented in the form of feature vectors. The image feature vectors can be directly used for model training. Take the optimized feature vectors as the feature dataset and store them in the vector database Milvus. Specifically, create a collection in the vector database Milvus to store the image feature vectors after feature dimension optimization in Step 4. The collection contains the following fields: image ID, image label, image path, and feature vectors, etc. Other fields can be added according to requirements. During model training, the feature vector data of the images can be directly obtained from Milvus for training.
[0065] Milvus is a high-performance and scalable vector database suitable for large-scale AI applications, especially in scenarios such as NLP (Natural Language Processing), computer vision, and recommendation systems. Milvus is compatible with PyTorch, TensorFlow, and Hugging Face, and is easy to integrate into the AI ecosystem. It supports distributed deployment, is suitable for cloud applications, and can be managed using Kubernetes. As a high-performance vector database, Milvus can efficiently store and retrieve large-scale image feature vectors and support fast similarity search. The image feature vectors retrieved from Milvus can be directly used for model training to improve training efficiency and model performance.
[0066] When training the target detection model subsequently, directly call the corresponding feature dataset from Milvus, and divide the feature dataset into a training set, a validation set, and a test set for training, validating, and testing the target detection model. Among them, the training set (Training Set) is used to train the model and adjust the model parameters to minimize the loss function. The validation set (Validation Set) is used to evaluate the model performance during training and adjust the hyperparameters. The test set (Test Set) is used to finally evaluate the generalization ability of the model and is usually used after the model training is completed.
[0067] The target detection model of this application can be a deep learning model such as a convolutional neural network (CNN) model, a recurrent neural network (RNN) model, and a Transformer model. Among them, the convolutional neural network model is used for image processing. The recurrent neural network model is used for processing sequence data (such as time series, text). The Transformer model is used for natural language processing and multimodal tasks. For example, for the Transformer model, its input is a sequence embedding in the form of [Batch_size, X, D]. Among them, Batch_size is the batch size, which refers to the number of data samples used in each iteration (or called "step") when training a machine learning or deep learning model. X is a variable-length sequence, and D is the embedding dimension, which includes the following two parts. Block content embedding: The features extracted from the image blocks, such as a 1×X-dimensional sequence, that is, the optimized feature vector generated in step 4. Position encoding: The position information added to each image block, such as the positions marked by [b] and [e]. [b] and [e] are used as delimiters to help the model identify the sequence boundaries.
[0068] The task requirements of the target detection model determine the output form of the model. For example, the output is sequence-to-sequence (Seq2Seq): output a sequence of the same length as the input, such as machine translation. Sequence-to-label (Seq2Label): output a single classification result, such as image classification. Autoregressive generation: generate subsequent tokens one by one, such as text generation. In the target detection task of this application, the following form is adopted.
[0069] Input: [b] + target block sequence + [e], with a shape of [Batch_size, X+2, D].
[0070] Output: Designed according to specific task requirements, for example: Predict the target category (classification task) → output [Batch_size, num_classes]. num_classes represents the number of target categories; Generate the target position (detection task) → output [Batch_size, X, 4], where 4 is the number of bounding box coordinates.
[0071] The trained object detection model can be used to implement the following functions. For example, object detection: locating the position of specific objects in an image. Image description generation: generating a natural language description based on a sequence of object blocks. Image restoration: reconstructing a complete image based on the remaining object blocks. Image-text retrieval: matching a sequence of image blocks with a text description.
[0072] The following lists several common application scenarios of the object detection model.
[0073] Application scenario 1: Medical image analysis - Tumor area localization. The task objective is to quickly locate the tumor area in CT / MRI images to assist doctors in diagnosis. The input of the trained object detection model is a sequence of medical image blocks preprocessed by the method of the present application, only retaining the object blocks classified as "tumor" (abbreviated as blocks), such as saved in the form of [b, block 1, block 2, ..., block X, e], and each block contains local tissue features. The output of the trained object detection model is a heat map of the tumor position (marking the coordinates of all image blocks classified as "1" in the original image) or a probability distribution map. The heat map or probability distribution map of the tumor position output by the model can be used to assist in diagnosis, reduce the doctor's image review time, and improve the efficiency of early cancer screening.
[0074] Application scenario 2: Autonomous driving - Real-time detection of road obstacles. The task objective is to identify obstacles such as pedestrians, vehicles, and objects on the road in front of the vehicle. The image or video stream captured by the on-vehicle camera is preprocessed by the method of the present application to generate a corresponding feature vector, that is, a sequence of object blocks. The input of the trained object detection model is a sequence of image blocks, only retaining the object blocks containing obstacles, such as saved in the form of [b, block A, block B, ..., block X, e]. The output of the trained object detection model is the coordinates of the obstacle bounding box (reconstructed based on the position of the image blocks) and the category (such as pedestrian, vehicle, object). Real-time detection of road obstacles based on the trained object detection model can effectively improve the real-time response ability of the autonomous driving system and avoid collisions.
[0075] Application scenario 3: Industrial quality inspection - Product defect detection. The task objective is to detect surface defects of products on the production line, such as cracks, scratches, bubbles, and spots. The image or video stream captured by the industrial camera on the production line is preprocessed by the method of the present application to generate a corresponding sequence of object blocks. The input of the trained object detection model is a sequence of image blocks of the product, only retaining the defect blocks, such as saved in the form of [b, defect block 1, ..., defect block X, e]. The output of the trained object detection model is the defect type classification (such as crack = 1, scratch = 2, bubble = 3, spot = 4) and the defect position marking. Detecting surface defects of products based on the trained object detection model can replace manual quality inspection and improve the detection efficiency and consistency.
[0076] Application Scenario 4, Security Monitoring: Abnormal Object Recognition. The task objective is to identify suspicious objects in surveillance videos, such as left-behind packages, dangerous goods, etc. The input of the trained object detection model is the sequence of target blocks after the video frames preprocessed by the method of this application are segmented, and only the abnormal blocks are retained, such as saved in the form of [b, abnormal block 1,..., abnormal block X, e]. The output of the trained object detection model is the alarm of the abnormal object category (such as "package", "knife") and the position coordinates. Based on the trained object detection model for abnormal object recognition, it can enhance the security of public places and reduce the cost of manual monitoring.
[0077] Application Scenario 5, Agricultural Monitoring: Identification of Pest and Disease Areas. The task objective is to locate the crop areas affected by pests and diseases through the farmland images taken by drones. The input of the trained object detection model is the sequence of farmland picture blocks preprocessed by the method of this application, and only the target blocks with pests and diseases are retained, called diseased blocks, such as saved in the form of [b, diseased block 1,..., diseased block X, e]. The output of the trained object detection model is the grading of the severity of pests and diseases (such as mild / severe) and the distribution heat map. Based on the trained object detection model for pest and disease area identification, it can assist in precise pesticide application, reduce pesticide waste, and increase crop yields.
[0078] The key points of the picture data preprocessing method of this application are as follows: on the one hand, the picture data is automatically screened according to the task objective and its target types, and the proportion of types is balanced; on the other hand, the dimensionality reduction of the picture data features is carried out. The feature engineering in the picture data preprocessing stage processes the picture data into feature vector data with a lower overall dimension. In this way, less hardware resources are required in the model training stage, that is, the space complexity is reduced. Under the same model and parameter conditions, the time complexity of training is also lower. At the same time, the processed feature vector data can be reused by the models composed of other algorithms, reducing resource consumption and development and maintenance costs, and improving the resource utilization efficiency and the efficiency of model development. Among them, the time complexity (Time Complexity) refers to the rate at which the running time of an algorithm grows with the increase of the input scale, usually represented by the big O notation, such as O(n) 2 ) etc. The space complexity (Space Complexity) refers to the size of the storage space occupied during the running of the algorithm, and is also represented by the big O notation.
[0079] To verify the effectiveness of this application, the feature vectors obtained by the method of this application are used to train the Transformer model. The advantages of this application in terms of time complexity and space complexity in the object detection task can be demonstrated through the following mathematical derivations and comparative analyses. According to the method of this application, the dimension of the feature vector after processing each image is 1×X, where X varies depending on the number of target blocks in the image. Although traditional neural networks (such as fully connected networks, CNNs) usually require fixed input sizes, the design of Transformer natively supports variable-length sequence inputs. The core of Transformer is the self-attention mechanism, and its computational process does not depend on the sequence length, so it can process input sequences of any length.
[0080] First, the computational complexity of Transformer mainly comes from the self-attention mechanism, and its time complexity and space complexity are both proportional to the square of the input sequence length n. The time complexity is expressed as O(n²d), where d is the feature dimension. The space complexity is expressed as O(n²), which is stored through the attention matrix. The following is a comparison of the input sequence lengths of the traditional method and the method of this application.
[0081] The traditional method uses the original pixel input. Assume that the size of the original image is h×w, and the input sequence length n1 = h×w. For example, when h = 224 and w = 224, n1 = 224×224 = 50,176.
[0082] For the method of this application, in the block division stage, the image is segmented into picture blocks of p × p pixels, and the number of blocks is (h / p) × (w / p). For example, when p = 16, the number of blocks = (224 / 16)×(224 / 16) = 14×14 = 196. In the block classification and recombination stage, only the target blocks marked as "1" are retained. Assume that k target blocks are retained (k < total number of blocks). If the target coverage area accounts for 10%, then k≈196×10% = 19.6≈20. In the sequence construction stage, special markers [b] and [e] are added, and the final sequence length n2 = k + m, where m is the number of special markers. Since the time and space complexity required for processing special markers are not in the same order of magnitude as that of target blocks, they can be ignored during calculation, that is, it is considered that n2 = k.
[0083] Then the comparison of time complexity is as follows.
[0084] The time complexity of the traditional method: T1 = O(n1²d) = O((h×w)²d). Numerical example: when h = 224, w = 224, and d = 768, T1≈50,176² × 768≈1.93×10¹² operations.
[0085] Time complexity of the method proposed in this application: T2 = O(n2²d) = O(k²d). Numerical example: When k = 20 and d = 768, T2 ≈ 20² × 768 operations.
[0086] Then the speedup ratio: T1 / T2 ≈ 1.93×10¹² / 3.07×10 5 ≈ 6.28×10 6 times.
[0087] The comparison of space complexity is as follows.
[0088] Space complexity of the traditional method: S1 = O(n1²) = O((h×w)²). Numerical example: S1 ≈ 50,176² ≈ 2.52×10 9 storage units.
[0089] Space complexity of the method proposed in this application: S2 = O(n2²) = O(k²). Numerical example: When k = 20, S2 ≈ 20² = 400 storage units.
[0090] Then the compression ratio: S1 / S2 ≈ 2.52×10 9 / 400 ≈ 6.3×10 6 times.
[0091] From the above comparison, it can be seen that the method of this application can achieve exponential compression of the sequence length. Through block division and classification filtering, the input sequence length is reduced from O(h×w) to O(k), and k << h×w. Mathematical relationship: k ≈ α×h×w / p², where α is the target coverage rate and α << 1. On the other hand, the method of this application can also achieve a quadratic decrease in time and space complexity. Since the complexity is related to n², compressing the sequence length can directly reduce the computational amount and storage requirements quadratically. In addition, the influence of the feature dimension d is that if the feature dimension d increases after block division (such as d = 768 in ViT), its influence is much smaller than the quadratic compression of the sequence length. Mathematical verification: Assume that d increases to 3 times, but n² decreases to 1 / 10 6 , the total complexity is still reduced to 1 / 3×10 6 . For the generalization of the actual scenario, the target coverage rate α is usually small (such as α < 20%), then the optimization effect of this application on the complexity will be very significant in most tasks.
[0092] The method of this application compresses the input sequence length from O(h×w) to O(k) (k << h×w) through block classification and sequence recombination, reducing the time complexity and space complexity of the Transformer model to 1 / α²×(p² / (h×w))² of the traditional method, and achieving a million-fold optimization under typical parameters. This feature is particularly applicable to high-resolution image processing tasks.
[0093] In an exemplary embodiment, this application also provides an image data preprocessing device, including: An image dataset acquisition module, configured to acquire a corresponding image dataset according to the task objective of the target detection model; the task objectives of the target detection model include human object detection, animal object detection, vehicle object detection, lesion area detection, obstacle detection, product defect detection, abnormal item detection, and pest and disease area detection; the corresponding image datasets include human image datasets, animal image datasets, vehicle image datasets, lesion image datasets, obstacle image datasets, product image datasets, item image datasets, and crop image datasets; An image classification and screening module, configured to classify and screen various image data in the image dataset according to the target types to be detected, and obtain a target type dataset; A data balancing processing module, configured to perform data balancing processing on the image data of each target type in the target type dataset to obtain a balanced dataset; A feature dimension optimization module, configured to perform feature dimension optimization processing on each image data in the balanced dataset to obtain an optimized feature vector; A feature vector storage module, configured to store the optimized feature vectors in a vector database for use in training the target detection model.
[0094] In an exemplary embodiment, this application also provides a computer device, which can be a server or a terminal. The computer device includes a processor, a memory, an input / output interface, and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, the image data preprocessing method described above is implemented.
[0095] In an exemplary embodiment, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-described picture data preprocessing method is implemented.
[0096] In an exemplary embodiment, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-described picture data preprocessing method is implemented.
[0097] Those of ordinary skill in the art can understand that all or part of the processes in the above-described embodiment methods can be completed by hardware related to computer program instructions. The computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it can include the processes of the above-described method embodiments. Among them, any reference to a memory or other medium provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0098] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0099] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0100] In this text, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. At the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for preprocessing picture data, characterized in that, Including: Obtain the corresponding image dataset according to the task objective of the object detection model; The task objectives of the object detection model include human object detection, animal object detection, vehicle object detection, lesion area detection, obstacle detection, product defect detection, abnormal item detection, and pest and disease area detection; the corresponding image datasets include human image dataset, animal image dataset, vehicle image dataset, lesion image dataset, obstacle image dataset, product image dataset, item image dataset, and crop image dataset; Classify and screen various image data in the image dataset according to the target types to be detected to obtain a target type dataset; Perform data balancing processing on the image data of each target type in the target type dataset to obtain a balanced dataset; Perform feature dimension optimization processing on each image data in the balanced dataset to obtain an optimized feature vector; Store the optimized feature vectors in a vector database for use in training the object detection model.
2. The picture data preprocessing method according to claim 1, characterized in that The step of classifying and screening various image data in the image dataset according to the target types to be detected to obtain a target type dataset specifically includes: Calculate the similarity between each image data in the image dataset and each target type through the CLIP model, and retain the image data with a similarity greater than the similarity threshold to form a target type dataset.
3. The picture data preprocessing method according to claim 2, wherein The step of performing data balancing processing on the image data of each target type in the target type dataset to obtain a balanced dataset specifically includes: Judge whether the proportion of image data between any two target types in the image data of each target type in the target type dataset is balanced; If the proportion is balanced, compress the image data of each target type to the same classification quantity Nmax; where Nmax is the number of images included in the target type with the largest proportion after data compression; If the proportion is unbalanced, perform data augmentation on the image data of the target type with a smaller proportion, perform data compression on the image data of the target type with a larger proportion, and enhance or compress the image data of each target type to the same classification quantity Nmax.
4. The picture data preprocessing method according to claim 3, wherein The step of performing feature dimension optimization processing on each image data in the balanced dataset to obtain an optimized feature vector specifically includes: Perform block processing on each image data in the balanced dataset, and divide each image data into multiple image blocks of M rows × N columns; both M and N are integers greater than 1; Perform block-level classification on each image block, mark the image block containing the target object as 1, and mark the image block not containing the target object as 0 to obtain labeled image blocks; Based on the labels carried by each image block, screen and reorganize the multiple image blocks into which each image data is divided to obtain an optimized feature vector.
5. The picture data preprocessing method according to claim 4, characterized in that The step of screening and reorganizing the multiple image blocks into which each image data is divided based on the labels carried by each image block to obtain an optimized feature vector specifically includes: Traverse the multiple image blocks of M rows × N columns into which each image data is divided, and take each row as a unit, retain the image blocks with a label of 1, and discard the image blocks with a label of 0; Recombine all picture blocks labeled 1 into a feature matrix with a dimension of 1×X. After each row of picture blocks is successfully recombined, add the sequence start position marker [b]. After the recombination of the last row of picture blocks is completed, add the sequence end position marker [e] to obtain an optimized feature vector; where X is the number of picture blocks labeled 1 in each picture data.
6. The picture data preprocessing method according to claim 5, characterized in that Storing the optimized feature vector in a vector database for use in training a target detection model specifically includes: Using the optimized feature vector as a feature data set and storing it in the vector database; When training a target detection model, divide the feature data set into a training set, a validation set, and a test set for training, validating, and testing the target detection model.
7. An image data preprocessing device, characterized in that, It includes: A picture data set acquisition module for acquiring a corresponding picture data set according to the task objective of the target detection model; The task objectives of the target detection model include human object detection, animal object detection, vehicle object detection, lesion area detection, obstacle detection, product defect detection, abnormal item detection, and pest and disease area detection; the corresponding picture data sets include human picture data sets, animal picture data sets, vehicle picture data sets, lesion picture data sets, obstacle picture data sets, product picture data sets, item picture data sets, and crop picture data sets; A picture classification and screening module for classifying and screening various picture data in the picture data set according to the target types to be detected to obtain a target type data set; A data balance processing module for performing data balance processing on the picture data of each target type in the target type data set to obtain a balanced data set; A feature dimension optimization module for performing feature dimension optimization processing on each picture data in the balanced data set to obtain an optimized feature vector; A feature vector storage module for storing the optimized feature vector in a vector database for use in training a target detection model.
8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the picture data preprocessing method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the picture data preprocessing method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the picture data preprocessing method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Certificate photo exposure direction detection algorithm based on grid feature extraction
CN107169507A
Pedestrian re-recognition method based on body decomposition and significance detection
CN108520226A
Unbalanced text classification method and device, equipment and storage medium
CN113869398A
High-density stress field prediction method and system based on T2T-ViT and storage medium
CN118153381A
Refined category image generation method based on retrieval enhancement
CN118170937A