Multi-modal data processing method and device, electronic equipment and storage medium
By processing multimodal data, including encoding and feature extraction of sample label text and vehicle images, the problems of low target detection and recognition efficiency and poor generalization ability in the prior art are solved, and more efficient and accurate target recognition model training is achieved.
Patent Information
- Application Number
- CN202411931049.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-06
AI Technical Summary
Existing object detection and recognition technologies are inefficient in specific scenarios and new categories, have poor generalization capabilities, and large-scale model training requires a large amount of computing resources and data.
By processing multimodal data, including encoding sample tag text, target recognition and feature extraction of vehicle images, filtering out areas containing target vehicles, and training the target recognition model based on these features.
It improves the training efficiency and accuracy of the target recognition model, reduces dependence on a large amount of training data and computing resources, and enhances the generalization ability of the model.
Smart Images

Figure CN119942471A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, electronic device and storage medium for processing multimodal data. Background Art
[0002] With the development of artificial intelligence technology, artificial intelligence (AI) technology has been applied to all aspects of human life. Target detection and recognition methods play a vital role in the field of artificial intelligence. This technology is widely used. Common target detection and recognition technologies are all developed based on convolutional neural networks (CNN).
[0003] The training of common target detection and recognition technologies usually requires a large amount of data, and the corresponding models need to be retrained for specific targets. The training of large models often consumes a lot of computing resources and requires a large amount of training data, which affects the efficiency of target detection and recognition for specific targets. For target detection and recognition tasks in specific scenarios, the classifier must be trained with training data from specific scenarios. When encountering new categories, the model often needs to be retrained, and the zero-sample learning ability is extremely poor and the generalization ability is not strong. Summary of the invention
[0004] In view of this, embodiments of the present application provide a method for processing multimodal data, an electronic device, a storage medium, and a program product.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The present application provides a method for processing multimodal data, the method comprising:
[0007] Encoding a plurality of sample label texts to obtain a first text feature of each of the sample label texts;
[0008] Performing target recognition processing on the first sample vehicle image to obtain a first predicted area in the first sample vehicle image;
[0009] Perform binary classification processing on the first prediction area to obtain a second prediction area containing the target vehicle in the first prediction area;
[0010] Performing feature extraction processing on the second prediction area to obtain target image features;
[0011] The target recognition model is trained based on the first text feature and the target image feature to obtain the trained target recognition model.
[0012] In the above solution, the encoding process is performed on the plurality of sample label texts to obtain the first text feature of each of the sample label texts, including:
[0013] For each of the sample label texts, the sample label text is expanded from words to sentences to obtain new sample label texts;
[0014] Each of the new sample label texts is encoded to obtain a first text feature of the sample label text.
[0015] In the above solution, the performing target recognition processing on the first sample vehicle image to obtain the first predicted area in the first sample vehicle image includes:
[0016] Performing feature extraction processing on the first sample vehicle image to obtain a first image feature;
[0017] Generate multiple candidate regions based on the first image feature, and determine the confidence of each candidate region, wherein the confidence is the probability that the candidate region contains the target vehicle;
[0018] At least one region with the highest confidence level is selected from the multiple candidate regions as the first prediction region.
[0019] In the above solution, the step of performing feature extraction processing on the first sample vehicle image to obtain the first image feature includes:
[0020] Performing convolution processing on the first sample vehicle image to obtain a first feature map;
[0021] Performing upsampling processing on the first feature map twice to obtain a second feature map and a third feature map, wherein the third feature map is obtained by performing the upsampling processing on the second feature map;
[0022] Performing a transposed convolution on the second feature map to obtain a fourth feature map;
[0023] The fourth feature map and the third feature map are fused to obtain the first image feature.
[0024] In the above scheme, the first image feature is represented in the form of a feature map;
[0025] The performing feature extraction processing on the second prediction area to obtain target image features includes:
[0026] Determining a position mapping relationship between the second prediction area and a feature map corresponding to the first image feature;
[0027] Based on the position mapping relationship, extracting a second image feature corresponding to the second prediction area from the first image feature;
[0028] Performing dimensionality reduction processing on the second image feature to obtain the target image feature, wherein the target image feature is represented as a feature map of a preconfigured size.
[0029] In the above scheme, the target recognition processing and the binary classification processing are implemented by a neural network model;
[0030] Before encoding the plurality of sample label texts to obtain the first text feature of each of the sample label texts, the method further includes:
[0031] Acquire a sample training set, wherein the sample training set includes: a plurality of second sample vehicle images, an actual position of a target vehicle area in each of the second sample vehicle images, and a label value, when a target vehicle exists in the target vehicle area, the label value is a first preset value, and when no target vehicle exists in the target vehicle area, the label value is a second preset value;
[0032] Based on the plurality of second sample vehicle images, calling the neural network model to perform target recognition processing to obtain a predicted position of a third predicted area in each of the second sample vehicle images;
[0033] Determining the regression loss of the neural network model according to the predicted position and the actual position;
[0034] Based on the multiple second sample vehicle images, calling the neural network model to perform binary classification processing to obtain a prediction confidence of each second sample vehicle image;
[0035] Determining the cross entropy loss of the neural network model according to the prediction confidence and the label value;
[0036] The neural network model is trained according to the cross entropy loss and the regression loss.
[0037] In the above solution, after training the target recognition model based on the first text feature and the target image feature to obtain the trained target recognition model, the method further includes:
[0038] Obtain an image to be recognized;
[0039] The trained target recognition model is called based on the image to be recognized to perform target recognition processing, and the position information of each target vehicle in the image to be recognized and the type text of each target vehicle are obtained.
[0040] The present application provides a multimodal data processing device, the device comprising:
[0041] A feature extraction module, used for encoding a plurality of sample label texts to obtain a first text feature of each of the sample label texts;
[0042] A target recognition module, used for performing target recognition processing on the first sample vehicle image to obtain a first predicted area in the first sample vehicle image;
[0043] The target recognition module is further used to perform binary classification processing on the first prediction area to obtain a second prediction area containing the target vehicle in the first prediction area;
[0044] The feature extraction module is further used to perform feature extraction processing on the second prediction area to obtain target image features;
[0045] A training model is used to train a target recognition model based on the first text feature and the target image feature to obtain the trained target recognition model.
[0046] An embodiment of the present application also provides an electronic device, comprising: a processor and a memory for storing a computer program that can be run on the processor, wherein the processor is used to execute the steps in the above-mentioned multimodal data processing method when running the computer program.
[0047] An embodiment of the present application further provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method for processing multimodal data are implemented.
[0048] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps in the above-mentioned multimodal data processing method.
[0049] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps in the above-mentioned multimodal data processing method.
[0050] The embodiment of the present application processes the sample vehicle image in two stages, specifically including predicting the first prediction area in the sample vehicle image where the target vehicle may exist, classifying the first prediction area, screening out the second prediction area containing the target vehicle, and training the target recognition model based on the target image features of the second prediction area, thereby improving the accuracy of obtaining samples containing the target vehicle, and thus improving the accuracy of the trained target recognition model. The label text is converted into text features, the sample vehicle image is converted into target image features, and the target recognition model is trained and processed through the text features and the target image features, so that the target recognition model is fine-tuned. The data used in the training process does not need to be manually labeled, and the target recognition model does not need to perform complex processing to screen out additional information, thereby improving the training efficiency of the target recognition model and the accuracy of the target recognition model obtained after training in identifying vehicles. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 A schematic diagram of the principle of a method for processing multimodal data provided in an embodiment of the present application;
[0052] Figure 2 A schematic diagram of the structure of the target recognition model provided in the embodiment of the present application;
[0053] Figure 3A A schematic diagram of a first flow chart of a method for processing multimodal data provided in an embodiment of the present application;
[0054] Figure 3B A second flow chart of the method for processing multimodal data provided in an embodiment of the present application;
[0055] Figure 3C A third flow chart of the method for processing multimodal data provided in an embodiment of the present application;
[0056] Figure 4 A schematic diagram of the text encoding principle provided in the embodiment of the present application;
[0057] Figure 5 A schematic diagram of the principles of picture encoding and detection frame prediction provided in an embodiment of the present application;
[0058] Figure 6 A schematic diagram of the principle of the training target recognition model provided in the embodiment of the present application;
[0059] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments.
[0061] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0062] With the development of artificial intelligence technology, artificial intelligence (AI) technology has been applied to all aspects of human life. Target detection and recognition methods play a vital role in the field of artificial intelligence. This technology is widely used. Common target detection and recognition technologies are all developed based on convolutional neural networks (CNN). The main methods are as follows: build a convolutional neural network model through convolutional layers, pooling layers, activation functions, etc., train detector classifiers, and complete target detection and recognition tasks. Common network structures include Visual Geometry Group deep learning models, residual networks (Resnet), Mobilenet, etc. The network structure is adjusted for different training tasks, and some mainstream detection and recognition methods are also included, such as Fast Region-based Convolutional Networks (Fast R-CNN), YOLO series, Single Shot MultiBox Detector (SSD), etc. In related technologies, large models used for target detection and recognition are usually trained in a self-supervised manner, but self-supervised training requires a large amount of data. The training of large models often requires a lot of computing resources, and retraining a large model for a specific goal takes a lot of time.
[0063] In the field of vehicle detection and recognition, most of the existing vehicle detection and recognition technologies use convolutional neural networks, which can complete the target detection and classification tasks of specific scenes. However, the classifier must use the training data of specific scenes to complete the model training. It often has poor results for new scenes that the model has not seen, and the zero-sample generalization ability is extremely poor. The target detection and recognition technology based on multi-source information fusion can complete the target detection and recognition tasks of specific scenes, but the classifier must use the training data of specific scenes to complete the training. When encountering new categories, it is often necessary to retrain the model, and the adaptability to new samples that have not been learned is not high. The data input of the target detection and recognition technology based on multi-source information fusion is all image data, which is a single modal data input and has poor generalization ability. The detection and recognition technology of the multimodal large model needs to train the models of the two modalities separately. The training of the detector still requires a large amount of target data to train in order to obtain a better detection model. Secondly, after the detector is trained, the multimodal large model can be fine-tuned, which is inefficient and cumbersome. Although the zero-sample learning ability is enhanced through the multimodal large model, the training of the detector still belongs to the traditional detection method, and the zero-sample learning ability is poor. Detection and recognition technologies based on multimodal large models often perform repeated extraction of image features. For example, features are extracted once for the entire image during detector training, and then the features need to be extracted again during fine-tuning of the large model. Repeated feature extraction affects the efficiency of the algorithm.
[0064] In view of this, an embodiment of the present application provides a method for processing multimodal data. The embodiment of the present application achieves end-to-end vehicle detection and recognition by fine-tuning a large multimodal model, encodes sample label text, and performs target recognition processing on a first sample vehicle image, segments an area containing a target vehicle from the first sample vehicle image, and trains a vehicle recognition model based on target image features of the area containing the target vehicle and encoded features of the sample label text, thereby improving the efficiency and training effect of target recognition model training, and thereby improving the accuracy of the target recognition model in detecting the target vehicle.
[0065] Figure 3A A first flow chart of a method for processing multimodal data provided in an embodiment of the present application is shown as follows: Figure 3A As shown, the method for processing multimodal data provided in the embodiment of the present application is as follows: Figure 3A Steps 301 to 305 are implemented as described in detail below.
[0066] In step 301, multiple sample label texts are encoded to obtain a first text feature of each sample label text.
[0067] For example, the sample label text is used to characterize the type of the sample image. Assuming that the sample image is a vehicle image, the sample label text may be the name of the vehicle type, such as a bus, a car, a motorcycle, and a bicycle. The encoding process may be implemented by a text encoder.
[0068] For example, to facilitate understanding of the multimodal data processing method provided in the embodiment of the present application, refer to Figure 1 , Figure 1 A schematic diagram of the principle of a method for processing multimodal data provided in an embodiment of the present application. In an embodiment of the present application, a sample label text (e.g., bus, car, motorcycle, etc.) is converted into a first text feature of the code by a text encoder 101, a sample image is subjected to feature extraction by a convolutional neural network 102 (CNN), a feature map is obtained, and the feature map is encoded based on a transformer model encoder 103 (Transformer Encoder), the encoded feature map is input into an upsampled and transposed convolution (Upsampled and Transposed Convolution, UPTC) module 104, a region proposal network (Region Proposal Network, RPN) 105 for processing in sequence, a predicted region in the feature map is extracted by a region of interest alignment (Region of Interest Align, ROI Align) module 106, and a predicted region is binary classified by a feed-forward neural network (Feed-Forward Neural Network) 107 to determine whether the predicted region type contains a target vehicle or does not contain a target vehicle, and the feed-forward neural network is also used to determine the coordinate information of the predicted region. The prediction area including the target vehicle and the coordinate information of the prediction area of the type are mapped to the feature map output by the converter model encoder 103, and the target image features corresponding to the prediction area including the target vehicle are obtained. The target recognition model is trained based on the first text features and the target image features to obtain a trained target recognition model. The target recognition model can be a contrastive language-image pre-training (CLIP) large model 110.
[0069] In some embodiments, step 301 can be implemented in the following manner: for each sample label text, the sample label text is expanded from words to sentences to obtain new sample label text; each new sample label text is encoded to obtain the first text feature of the sample label text.
[0070] For example, the sample label text can be expanded from a word to a sentence by the following method: fill the sample label text into the empty space of the preset sentence template to obtain a new sample label text. For example: the labels of 5 types of vehicles, bus, car, bicycle, truck, and motor, are summarized in a sentence as a photo of a bus, a photo of a car, a photo of a bicycle, a photo of a truck, and a photo of a car (a photo of bus, a photo of car, a photo of bicycle, a photo of truck, a photo of motor), and the sentence template is "a photo of {}", and {} is used to fill in the name of the object, and the word label is converted into a sentence. This is because most of the training data of the CLIP model in the training process is a sentence, and converting the word label into a sentence can effectively improve the classification accuracy. During the training process of the CLIP model, the converted sentence will be converted into the feature vector of each sentence by a text encoder text Encoder. In the embodiment of the present application, the text encoder (text Encoder) can use a text encoder such as (Bidirectional Encoder Representations from Transformers, BERT). Figure 4 A schematic diagram of the principle of text encoding provided in an embodiment of the present application, in which labels such as bus, car, bicycle, truck, and automobile are converted into sentences, and the sentences are input into the text encoder 101 to obtain the first text feature of each label (for example: T1, T2, etc.).
[0071] In step 302, target recognition processing is performed on the first sample vehicle image to obtain a first predicted area in the first sample vehicle image.
[0072] For example, the target recognition process can be implemented in the following manner: feature extraction is performed on the first sample vehicle image to obtain a feature map, and a first prediction region that may contain the target vehicle is predicted based on the feature map. The feature extraction process is implemented by a convolutional neural network, and the target recognition process can be implemented by a region proposal network (RPN), which is a method for generating a pre-selected box in a fast region convolutional neural network (faster RCNN).
[0073] In some embodiments, reference Figure 3B , Figure 3B This is a second flow chart of the method for processing multimodal data provided in an embodiment of the present application. Step 302 can be implemented through steps 3021 to 3023, which are described in detail below.
[0074] In step 3021, feature extraction processing is performed on the first sample vehicle image to obtain a first image feature.
[0075] For example, the feature extraction process can be implemented by a convolutional neural network and an encoder.
[0076] In some embodiments, step 3021 can be implemented by: performing convolution processing on the first sample vehicle image to obtain a first feature map; performing upsampling processing on the first feature map twice to obtain a second feature map and a third feature map, wherein the third feature map is obtained by performing upsampling processing based on the second feature map; performing transposed convolution on the second feature map to obtain a fourth feature map; and performing fusion processing on the fourth feature map and the third feature map to obtain the first image feature.
[0077] Figure 5 Schematic diagram of the principle of picture encoding and detection frame prediction provided in the embodiment of the present application; the first sample vehicle image is passed through a shallow convolutional neural network 102 (CNN) to output a feature map. During the picture encoding process, a shallow convolutional neural network is first used and then input into the transformer model (transformer) encoder 103 to obtain a first feature map. The convolutional neural network pays more attention to local features, and the attention mechanism layer (Attention) in the transformer model pays more attention to global features. The use of a convolutional neural network combined with a transformer model can enhance the feature expression capability. The original image can be scaled by a convolutional neural network. Because the feature map after convolution will be reduced, the feature map finally input to the transformer model will become smaller, which can greatly reduce the amount of calculation of the transformer model. Before being input into the transformer model, the feature map is reduced in dimension through a convolutional layer CONV (1*1). The purpose of the dimensionality reduction is to reduce the amount of input parameters of the transformer model. The encoder of the transformer model converts the reduced feature map into a feature vector form.
[0078] The feature vector of the feature map after passing through the converter model encoder is flattened and input into an upsampled and transposed convolution (UPTC) network module 104, which mainly expands the feature vector of the spliced feature map through two upsampling and one deconvolution. Specifically, the size of the flattened feature map (first feature map) is F (c, h, w). After inputting the UPTC module, the second feature map of size F1 (c, 2h, 2w) is obtained through the first upsampling, and the third feature map of size F2 (c, 4h, 4w) is obtained through the second upsampling. F1 (c, 2h, 2w) is input into the 2x2 deconvolution with stride = 2 to obtain the fourth feature map FT2 (c, 4h, 4w), and finally F2 (c, 4h, 4w) and FT2 (c, 4h, 4w) are feature fused to obtain the first image feature Fout = F2 + FT2.
[0079] In step 3022, multiple candidate regions are generated based on the first image feature, and the confidence of each candidate region is determined.
[0080] For example, after obtaining the first image feature, a target pre-selection box is generated through a region selection network (RPN) 106, and the confidence of a valid pre-selection box (at least one) and the coordinate information of the pre-selection box are obtained through the RPN network. The confidence is the probability that the candidate region contains the target vehicle.
[0081] The region selection network uses a sliding window technique on the feature map, and each window center corresponds to a point on the feature map. For each window center, multiple anchor boxes of predefined sizes and proportions are generated. These anchor boxes are a hypothesis of the shape and size that the actual object may appear. The area corresponding to the anchor box is also the predicted area in the embodiment of the present application.
[0082] For each anchor box, the RPN network predicts two parameters: confidence (Classification Scores) and location information (Bounding Box Regression). The network outputs two confidence scores, one representing the probability that the anchor box contains an object, and the other representing the probability that the anchor box is the background. This is usually achieved through a classification layer (such as a softmax layer). The network also predicts four coordinate values, which represent the offset between the actual bounding box and the anchor box. These offsets are used to adjust the anchor box to better match the location of the actual object.
[0083] In step 3023, at least one region with the highest confidence is selected from the multiple candidate regions as the first prediction region.
[0084] For example, the confidence score reflects the network's judgment on whether the anchor box contains an object, while the position information is used to fine-tune the position of the anchor box so that it can more accurately frame the object in the image. Combining the two, a set of pre-selected boxes with confidence and precise positions can be obtained. Non-maximum Suppression (NMS), in order to remove redundant pre-selected boxes, non-maximum suppression technology is usually applied. Non-maximum suppression selects the pre-selected box with the highest confidence, and then removes those pre-selected boxes that overlap too much with it, thereby obtaining a set of high-confidence final pre-selected boxes.
[0085] In step 303, a binary classification process is performed on the first prediction area to obtain a second prediction area containing the target vehicle in the first prediction area.
[0086] For example, continue to refer to Figure 5 The prediction region output by the region selection network is extracted with features of related regions through the Region of Interest Align (ROIAlign) module 106, and the features of these ROI regions are input into the full connection. Based on the output result of the full connection layer, the feedforward neural network 107 performs a binary classification process of determining the second prediction region in the first prediction region and whether the position of the second prediction region belongs to the foreground or the background.
[0087] For example, foreground refers to the area of interest in an image, which is usually the focus of image analysis. In many cases, the foreground contains the main or target object in the image, such as a person, a vehicle, or any object that requires special attention. In image segmentation, the foreground is often the part separated from the background and needs further processing or analysis. The background refers to the area of the image that does not contain the main information or analytical interest. The background usually includes the environment behind the foreground object in the image or any part that does not require special attention. In image segmentation tasks, the background is often used to distinguish it from the foreground so that the foreground can be processed and analyzed more easily.
[0088] In step 304, feature extraction processing is performed on the second prediction area to obtain target image features.
[0089] In some embodiments, the first image feature is represented in the form of a feature map; step 304 can be achieved by: determining a position mapping relationship between the second prediction region and the feature map corresponding to the first image feature; based on the position mapping relationship, extracting the second image feature corresponding to the second prediction region from the first image feature; performing dimensionality reduction processing on the second image feature to obtain a target image feature, wherein the target image feature is represented as a feature map of a preconfigured size.
[0090] The feature map corresponding to the first image feature is that after the first feature map is input into the UPTC module, the second feature map of size F1(c, 2h, 2w) is obtained through the first upsampling, and the third feature map of size F2(c, 4h, 4w) is obtained through the second upsampling. F1(c, 2h, 2w) is input into the 2x2 deconvolution with stride=2 to obtain the fourth feature map FT2(c, 4h, 4w), and finally F2(c, 4h, 4w) and FT2(c, 4h, 4w) are feature fused to obtain the first image feature Fout=F2+FT2. The features of the second predicted area are mapped to the area corresponding to the first image feature according to the position mapping relationship, and the regional position of the second predicted area in the first image feature is obtained. The features in the regional position are extracted to obtain the second image feature of the second image feature. Reference Figure 6 , Figure 6 Schematic diagram of the principle of the training target recognition model provided for the embodiment of the present application. The first image feature output by the multiplexing converter model encoder 103 is used to map the predicted area position of the second predicted area predicted as the foreground to the first image feature, and obtain the position information (box) of the predicted area position on the first image feature. For example, if the output feature size of the second prediction area is 1 / 16 (h*w) of the input image, then the predicted area position information of the second prediction area is divided by 16 to obtain the position information of the ROI on the feature map of the first image feature, and input it into the region of interest alignment module 108 to obtain the second image feature. The second image feature is passed through a fully connected layer 109 to finally output the target image feature of the image. Since the fully connected layer converts the multi-dimensional feature map of the previous layer into a vector of a fixed size, it can be regarded as a dimensionality reduction process, reducing the amount of data for subsequent processing.
[0091] In step 305, the target recognition model is trained based on the first text feature and the target image feature to obtain a trained target recognition model.
[0092] For example, the first text feature and the target image feature are respectively input into the CLIP large model network for comparative learning, and the fine-tuned trained CLIP large model network is used to output specific category information of each target in the image. The CLIP large model network is used to determine the similarity between the text encoding and the image encoding, and the label corresponding to the text encoding with a similarity greater than a threshold is used as the specific category of the object in the target frame in the image encoding.
[0093] In the embodiment of the present application, an end-to-end vehicle detection and recognition architecture with fine-tuning of a multimodal large model is proposed. An end-to-end detection and recognition method based on fine-tuning of a CLIP multimodal large model is adopted. Text information and image information are input, and the text information is encoded at the same time. The shared converter model is used to extract features for the image information, and target box prediction is performed and feature encoding is performed on the predicted target area. Then, fine-tuning learning is performed through image-text comparison. On the basis of outputting the image category, the detection task can also be completed to achieve end-to-end target detection and recognition. In the process of image processing, the converter model encoder is shared. After encoding, in order to prevent the feature map from being too small, the UPTC module is used to expand the feature map through simple upsampling and deconvolution feature fusion, and then it is divided into two paths. One branch performs target box regression through RPN+ROIAlign, where RPN is a candidate box generation network, which is derived from fasterRCNN and is mainly used to generate candidate boxes on feature maps, and then extract features in the candidate boxes through ROIAlign, and then predict the precise position of the target through binary classification and regression. The other branch only extracts the ROI feature area predicted as the foreground by the previous branch through binary classification and regression, and encodes the feature through ROIAlign and the fully connected layer, and compares it with the text encoding information for learning. The shared module can effectively reduce the repeated extraction of target features, and only extract features through the converter model encoder once.
[0094] In some embodiments, before step 301, reference is made to Figure 3C , Figure 3C The third flow chart of the method for processing multimodal data provided in the embodiment of the present application is as follows: Steps 3061 to 3065 are executed to train the region selection network, which is described in detail below.
[0095] In step 3061, a sample training set is obtained.
[0096] Here, the sample training set includes: multiple second sample vehicle images, the actual position of the target vehicle area in each second sample vehicle image, and a label value. When there is a target vehicle in the target vehicle area, the label value is a first preset value, and when there is no target vehicle in the target vehicle area, the label value is a second preset value.
[0097] For example, when there is a target vehicle in the target vehicle area, the label value is 1, and when there is no target vehicle in the target vehicle area, the label value is 0. The actual position of the target vehicle area can be represented as t i * , representing the true coordinate values of the corner points of the target vehicle area.
[0098] In step 3062, based on the multiple second sample vehicle images, the neural network model is called to perform target recognition processing to obtain the predicted position of the third prediction area in each second sample vehicle image.
[0099] For example, the principle of target recognition processing can refer to the above steps 303 to 304, which will not be repeated here. The predicted position obtained by the region selection network prediction can be represented as t i =(t x , t y , t w , t h ), t i is a coordinate vector, representing the four parameter coordinates of the predicted coordinate frame regression t x , t y , t w , t h .
[0100] In step 3063, the regression loss of the neural network model is determined based on the predicted position and the actual position.
[0101] For example, the regression loss is represented by the following formula (1):
[0102] L reg (t i , t i * )=smooth L1 (t i -t i * ) (1)
[0103] Among them, the coordinate regression loss L reg The smoothL1 loss is used, which is a combination of L1 and L2 losses. L1 loss is not differentiable at 0, and the gradient of L2 loss is prone to explosion when the predicted value is very different from the target value. smoothL1 is represented by the following formula (2):
[0104]
[0105] The smoothL1 loss improves the shortcomings of both and combines the two losses using a piecewise function. For each target box, L is calculated. reg (t i , t i * ) After that, the coordinate regression loss is combined with the label value p i * When the object is present (positive sample), the label value is 1, and when there is no object (negative sample), the label value is 0, which means that only the foreground is calculated for loss, and the background is not calculated for loss. reg is the size of the feature map.
[0106] In step 3064, based on the plurality of second sample vehicle images, the neural network model is called to perform binary classification processing to obtain the prediction confidence of each second sample vehicle image. The cross entropy loss of the neural network model is determined according to the prediction confidence and the label value.
[0107] For example, the principle of binary classification can be referred to above, and will not be repeated here. Classification loss L cls Using the cross entropy loss function, the classification loss is calculated as the average loss of all samples, N cls The parameter is the total number of samples.
[0108] In step 3065, the neural network model is trained according to the cross entropy loss and the regression loss.
[0109] For example, the total loss L corresponding to the cross entropy loss and the regression loss is represented by formula (3):
[0110]
[0111] Based on the total loss, the region selection network is back-propagated and the parameters of the region selection network are updated to reduce the total loss, thereby realizing the training process for the region selection network.
[0112] In some embodiments, after step 305, the following processing is performed: obtaining the image to be identified; calling the trained target recognition model to perform target recognition processing based on the image to be identified, and obtaining the location information of each target vehicle in the image to be identified, as well as the type text of each target vehicle.
[0113] For example, the target recognition model encodes the image to be recognized to obtain the encoded features, determines the similarity between the encoded features and the text features of the label, and uses the label corresponding to the text feature with the greatest similarity as the type of the target vehicle in the image. At the same time, the target recognition model also determines the location information of the area where the target vehicle is located by aligning the region of interest.
[0114] In some embodiments, the multimodal data processing method provided in the embodiments of the present application can be applied in the field of vehicle identification technology. For example, by using cameras installed on the road to photograph accident-prone sections, the CLIP model trained based on the multimodal data processing method provided in the embodiments of the present application performs target recognition on the photographed photos, obtains the various vehicles involved in the image, and then analyzes the causes of traffic accidents.
[0115] In the embodiment of the present application, the target detection and recognition method based on multi-source information fusion, the target detection and recognition technology of multi-source information fusion is a new technology that realizes accurate target recognition by jointly processing data from multiple information sources and extracting the fusion features of the target. This technology can solve the problem of insufficient precision in target detection and recognition applications. Although multi-source information is fused in the related art, the fused features are all picture features, and multimodal features are not involved. The network structure used is also the traditional CNN structure. Compared with the multimodal method, this type of method only uses picture features, and the zero-sample learning ability is poor. Traditional deep learning model training often focuses on one problem. Take the detection model as an example. If a specific object is to be detected, the model needs to be trained based on the image of the specific object. However, if the model needs to detect something else, a new picture is needed to train the model and repeat the training. The multimodal large model has zero-shot capability and performs very well in tasks such as text-image retrieval, image classification, and image generation based on text. Compared with the current target detection and classification algorithm, the embodiment of the present application constructs a binary classification and bounding box regression network and a multimodal large model fine-tuning network, sharing a converter model encoder, reducing repeated feature extraction, saving computing resources, and realizing end-to-end vehicle detection and recognition functions, which is simple and efficient.
[0116] The embodiment of the present application also provides a device for processing multimodal data, which corresponds to the above-mentioned method for processing multimodal data. The steps in the embodiment of the method for processing multimodal data are also fully applicable to the embodiment of the present device.
[0117] The device comprises: a feature extraction module, which is used to encode a plurality of sample label texts to obtain a first text feature of each sample label text;
[0118] A target recognition module, used for performing target recognition processing on the first sample vehicle image to obtain a first predicted area in the first sample vehicle image;
[0119] The target recognition module is further used to perform binary classification processing on the first prediction area to obtain a second prediction area containing the target vehicle in the first prediction area;
[0120] The feature extraction module is also used to perform feature extraction processing on the second prediction area to obtain target image features;
[0121] The training model is used to train the target recognition model based on the first text feature and the target image feature to obtain a trained target recognition model.
[0122] It should be noted that: the multimodal data processing device provided in the above embodiment only uses the division of the above program modules as an example when processing multimodal data. In actual applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the device is divided into different program modules to complete all or part of the processing described above. In addition, the multimodal data processing device provided in the above embodiment and the multimodal data processing method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0123] Based on the hardware implementation of the above program modules and in order to implement the method of the embodiment of the present application, the embodiment of the present application also provides an electronic device. Figure 7 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the electronic device includes:
[0124] Communication interface 701, capable of exchanging information with other devices such as network devices;
[0125] The processor 702 is connected to the communication interface 701 to implement information exchange with other devices and is used to execute the method provided by one or more technical solutions when running a computer program. The computer program is stored in the memory 703.
[0126] Of course, in actual application, the various components in the electronic device are coupled together through the bus system 704. It can be understood that the bus system 704 is used to realize the connection and communication between these components. In addition to the data bus, the bus system also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 7 Various buses are labeled as bus system 704 .
[0127] The memory 703 in the embodiment of the present application is used to store various types of data to support the operation of the computer device. Examples of such data include: any computer program used to operate on the electronic device.
[0128] It can be understood that the memory 703 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAM bus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.
[0129] The method disclosed in the above embodiment of the present application can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present application, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in a memory, and the processor reads the program in the memory and completes the steps of the above method in combination with its hardware.
[0130] Optionally, when the processor 702 executes the program, it implements the corresponding processes implemented by the electronic device in each method of the embodiments of the present application, which will not be described in detail here for the sake of brevity.
[0131] In an exemplary embodiment, the present application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, for example, including a first memory storing a computer program, and the computer program can be executed by a processor of a computer device to complete the steps of the aforementioned method. The computer-readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface storage, optical disk, or CD-ROM.
[0132] In the several embodiments provided in the present application, it should be understood that the disclosed devices, computer equipment and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0133] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0134] In addition, all functional units in the embodiments of the present application may be integrated into one processing unit, or each unit may be a separate unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0135] A person of ordinary skill in the art can understand that: all or part of the steps of implementing the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium, which, when executed, executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, disks or optical disks.
[0136] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiment of the present application can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
[0137] In an exemplary embodiment, the embodiment of the present application further provides a computer program product, including a computer program, which can be executed by the processor 702 of the electronic device to complete the steps described in the method of the embodiment of the present application applied to the remote management terminal.
[0138] In an exemplary embodiment, the embodiment of the present application further provides a computer program product, including a computer program, which can be executed by the processor 702 of the electronic device to complete the steps described in the method applied to the terminal in the embodiment of the present application.
[0139] It should be noted that: "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0140] In addition, the technical solutions described in the embodiments of the present application can be combined arbitrarily without conflict.
[0141] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A method for processing multimodal data, characterized in that: The method comprises: Encoding a plurality of sample label texts to obtain a first text feature of each of the sample label texts; Performing target recognition processing on the first sample vehicle image to obtain a first predicted area in the first sample vehicle image; Perform binary classification processing on the first prediction area to obtain a second prediction area containing the target vehicle in the first prediction area; performing feature extraction processing on the second prediction area to obtain target image features; The target recognition model is trained based on the first text feature and the target image feature to obtain the trained target recognition model.
2. The method according to claim 1, characterized in that: The encoding process is performed on the plurality of sample label texts to obtain the first text feature of each of the sample label texts, including: For each of the sample label texts, the sample label text is expanded from words to sentences to obtain new sample label texts; Each of the new sample label texts is encoded to obtain a first text feature of the sample label text.
3. The method according to claim 1, characterized in that The performing target recognition processing on the first sample vehicle image to obtain a first predicted area in the first sample vehicle image includes: Performing feature extraction processing on the first sample vehicle image to obtain a first image feature; Generate multiple candidate regions based on the first image feature, and determine the confidence of each candidate region, wherein the confidence is the probability that the candidate region contains the target vehicle; At least one region with the highest confidence level is selected from the multiple candidate regions as the first prediction region.
4. The method according to claim 3, characterized in that The step of performing feature extraction processing on the first sample vehicle image to obtain a first image feature includes: Performing convolution processing on the first sample vehicle image to obtain a first feature map; Performing upsampling processing on the first feature map twice to obtain a second feature map and a third feature map, wherein the third feature map is obtained by performing the upsampling processing on the second feature map; Performing a transposed convolution on the second feature map to obtain a fourth feature map; The fourth feature map and the third feature map are fused to obtain the first image feature.
5. The method according to claim 3, characterized in that: The first image feature is represented in the form of a feature map; The performing feature extraction processing on the second prediction area to obtain target image features includes: Determining a position mapping relationship between the second prediction area and a feature map corresponding to the first image feature; Based on the position mapping relationship, extracting a second image feature corresponding to the second prediction area from the first image feature; Performing dimensionality reduction processing on the second image feature to obtain the target image feature, wherein the target image feature is represented as a feature map of a preconfigured size.
6. The method according to any one of claims 1 to 5, characterized in that: The target recognition process and the binary classification process are implemented by a neural network model; Before encoding the plurality of sample label texts to obtain the first text feature of each of the sample label texts, the method further includes: Acquire a sample training set, wherein the sample training set includes: a plurality of second sample vehicle images, an actual position of a target vehicle area in each of the second sample vehicle images, and a label value, when a target vehicle exists in the target vehicle area, the label value is a first preset value, and when no target vehicle exists in the target vehicle area, the label value is a second preset value; Based on the plurality of second sample vehicle images, calling the neural network model to perform target recognition processing to obtain a predicted position of a third predicted area in each of the second sample vehicle images; Determining the regression loss of the neural network model according to the predicted position and the actual position; Based on the multiple second sample vehicle images, calling the neural network model to perform binary classification processing to obtain a prediction confidence of each second sample vehicle image; Determining the cross entropy loss of the neural network model according to the prediction confidence and the label value; The neural network model is trained according to the cross entropy loss and the regression loss.
7. The method according to any one of claims 1 to 5, characterized in that: After training the target recognition model based on the first text feature and the target image feature to obtain the trained target recognition model, the method further includes: Obtain an image to be recognized; The trained target recognition model is called based on the image to be recognized to perform target recognition processing, and the position information of each target vehicle in the image to be recognized and the type text of each target vehicle are obtained.
8. A multimodal data processing device, characterized in that: The device comprises: A feature extraction module, used for encoding a plurality of sample label texts to obtain a first text feature of each of the sample label texts; A target recognition module, used for performing target recognition processing on the first sample vehicle image to obtain a first predicted area in the first sample vehicle image; The target recognition module is further used to perform binary classification processing on the first prediction area to obtain a second prediction area containing the target vehicle in the first prediction area; The feature extraction module is further used to perform feature extraction processing on the second prediction area to obtain target image features; A training model is used to train a target recognition model based on the first text feature and the target image feature to obtain the trained target recognition model.
9. An electronic device, characterized in that: include: A processor and a memory for storing a computer program that can be executed on the processor, wherein: The processor is used to execute the steps of the multimodal data processing method according to any one of claims 1 to 7 when running a computer program.
10. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for processing multimodal data according to any one of claims 1 to 7 are implemented.