Vehicle identification and positioning method and system based on deep learning visual positioning model

By adopting a multimodal fusion model based on deep learning in an intelligent transportation system, combining image features and natural language description for vehicle recognition and positioning, the problem of degradation of recognition efficiency and accuracy in complex traffic scenarios is solved, and more efficient and flexible vehicle management is achieved.

CN119919630APending Publication Date: 2025-05-02WUHAN YANGTZE COMM ZHILIAN TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411943491.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The prior art is difficult to accurately locate vehicles in complex traffic scenarios by combining natural language descriptions and image features, resulting in a decrease in recognition efficiency and accuracy, and cannot meet the needs of personalized target positioning and efficient vehicle management in intelligent traffic systems.

Method used

Using a visual positioning model based on deep learning, a multimodal fusion model is constructed, and the image features and text features in the traffic image dataset are trained to achieve vehicle recognition and positioning. The model includes data preprocessing, text feature coding, visual feature extraction, and multimodal feature fusion, which can be dynamically adjusted to match the user's natural language description.

Benefits of technology

It significantly improves the accuracy and efficiency of vehicle identification in complex scenarios, improves the interpretability of the model, and enables intelligent transportation systems to deal with diverse vehicle types and dynamic traffic scenarios more flexibly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919630A_ABST
    Figure CN119919630A_ABST
Patent Text Reader

Abstract

The invention provides a vehicle identification and positioning method and system based on a deep learning visual positioning model, and the method comprises the steps: collecting a traffic image, carrying out the data set pre-operation of the traffic image, and constructing a training data set; acquiring a vehicle driving real-time image, and preprocessing the vehicle driving real-time image to obtain a preprocessed real-time image; constructing a multi-modal fusion model, and training and adjusting the multi-modal fusion model by using the image features and the text features in the training data set to obtain a vehicle identification and positioning model; and inputting the preprocessed real-time image into the vehicle identification and positioning model, and outputting a vehicle identification and positioning result. According to the invention, through combination of deep learning and visual positioning technologies, accurate positioning and segmentation of the target vehicle are realized, and a new technical idea is provided for development of an intelligent traffic system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning technology, and in particular to a vehicle identification and positioning method and system based on a deep learning visual positioning model. Background Art

[0002] In the field of intelligent transportation, vehicle recognition through input text description is an emerging task. Its core is to quickly locate and mark the corresponding target vehicle in the traffic image based on the natural language description provided by the user. For example, when the input is "a red truck near the roadside", the system can combine the semantic information and image features in the text description to accurately find the location of the target vehicle in the traffic scene and mark it out. This technology is particularly suitable for complex traffic scenes, such as mixed vehicles of multiple types, different perspectives or partial occlusion of vehicles. It effectively improves the recognition efficiency and accuracy of target vehicles and provides new solutions for application scenarios such as traffic monitoring, violation detection and vehicle management.

[0003] The existing vehicle recognition scheme is mainly based on the traditional pure computer vision image processing algorithm, including vehicle detection, feature extraction and classification. For example, the image quality is usually improved by image preprocessing techniques (such as denoising, enhancement and standardization), and then the vehicle detection algorithm (such as background difference method, frame difference method, optical flow method, etc.) is used to extract the vehicle candidate area. Subsequently, feature extraction algorithms (such as HOG, SIFT, etc.) are used to generate feature vectors, and classification algorithms (such as SVM, decision tree, etc.) are used to classify the vehicle type and location. However, this method faces many limitations in practical applications. First, the traditional method is highly dependent on manually designed features, and it is difficult to adapt to the mixed, variable angles, changing lighting conditions and occlusion of multiple types of vehicles in complex traffic scenes, resulting in a significant decrease in the accuracy of detection and classification. Secondly, the traditional method lacks the ability to understand semantic information and cannot be dynamically adjusted according to the user's natural language description. For example, when it is necessary to identify the "red truck in the left lane" or the "blue car parked on the side of the road", the traditional method is unable to handle such a task because its core is still based on a single visual feature extraction and classification. Finally, these methods usually lack sufficient flexibility and adaptability in diverse and real-time demanding scenarios, resulting in limited application scope.

[0004] Existing technologies mainly rely on image features for vehicle identification, but when faced with diverse and complex traffic scenarios, there is a lack of an effective mechanism that combines natural language descriptions and image features to accurately locate vehicles. For example, when it is necessary to quickly extract the target vehicle through a specific description (such as "the white sedan in the left lane"), existing technologies cannot flexibly understand the semantic information in the text and efficiently match it with image features. This limitation significantly reduces the adaptability of the technology in practical applications, making it difficult to meet the needs of personalized target positioning and efficient vehicle management in intelligent transportation systems, limiting its wide application in scenarios such as real-time monitoring, dynamic management, and specific target search. Summary of the invention

[0005] The present invention provides a vehicle identification and positioning method and system based on a deep learning visual positioning model, so as to solve the defects existing in the prior art.

[0006] In a first aspect, the present invention provides a vehicle identification and positioning method based on a deep learning visual positioning model, comprising: Collecting traffic images, performing data set pre-operation on the traffic images, and constructing a training data set; Acquire a real-time image of a vehicle traveling, and preprocess the real-time image of the vehicle traveling to obtain a preprocessed real-time image; Constructing a multimodal fusion model, and using the image features and text features in the training data set to train and adjust the multimodal fusion model to obtain a vehicle recognition and positioning model; The pre-processed real-time image is input into the vehicle recognition and positioning model, and a vehicle recognition and positioning result is output.

[0007] According to a vehicle identification and positioning method based on a deep learning visual positioning model provided by the present invention, a traffic image is collected, a data set pre-operation is performed on the traffic image, and a training data set is constructed, including: Collect several traffic images; Performing fuzzy processing on the license plate in each traffic image to obtain a fuzzy processed image; Manually draw the outline of each vehicle on each blurred image to obtain a vehicle outline image; defining a keyword set, the keyword set comprising color, vehicle type, location information, and urban infrastructure; Write a text description for each vehicle based on the words in the keyword set; Combine the processed images with the text description of each vehicle to construct a training dataset.

[0008] According to a vehicle identification and positioning method based on a deep learning visual positioning model provided by the present invention, a real-time image of a vehicle driving is acquired, and the real-time image of the vehicle driving is preprocessed to obtain a preprocessed real-time image, including: Deploy camera modules at preset traffic monitoring points to collect real-time image and video data of vehicles in motion; Gaussian filtering is used to remove Gaussian noise in the image, median filtering is used to calculate the median of each pixel area to remove salt and pepper noise in the image, and bilateral filtering is used to maintain image edge information and smooth the interior of the image area to obtain a denoised image; Adaptive histogram equalization and improved adaptive histogram equalization are used to limit the contrast enhancement of the image, and gamma correction is used to adjust the image brightness to obtain the enhanced image; Mean-variance normalization is used to normalize the image pixel values ​​to a distribution with a mean of 0 and a variance of 1. Min-Max normalization is used to normalize the image pixel values ​​to between 0 and 1 or other specified intervals. Color channel normalization is used to normalize the RGB channels of the image separately to balance the distribution of each channel to obtain a standardized image.

[0009] According to a vehicle identification and positioning method based on a deep learning visual positioning model provided by the present invention, a real-time image of a vehicle driving is acquired, and the real-time image of the vehicle driving is preprocessed to obtain a preprocessed real-time image, and further includes: Use MySQL to build a traffic image database; An image acquisition module is established to acquire the pre-processed real-time image.

[0010] According to a vehicle identification and positioning method based on a deep learning visual positioning model provided by the present invention, a multimodal fusion model is constructed, and the multimodal fusion model is trained and adjusted using the image features and text features in the training data set to obtain a vehicle identification and positioning model, including: Establish a text feature encoding module, the text feature encoding module includes a single-layer bidirectional gated recurrent unit, input the natural language description into the text feature encoding module, convert it into a high-dimensional semantic feature vector, use the text feature encoding module as a language encoder, splice the bidirectional hidden state in each time step, generate word features, and provide language information for subsequent fusion operations; In the visual encoder, multi-scale features are down-sampled and flattened from the finest to the coarsest spatial resolution to generate visual features, providing image information for subsequent fusion operations; In order to achieve the fusion of text features and visual features, word features and visual features are input into the fusion module, language features are constructed through maximum pooling of channel dimensions, and Hadamard product operations are performed on word features and visual features to generate multimodal features. The multimodal features are input into the Transformer encoder to obtain the final vehicle recognition and positioning model.

[0011] According to a vehicle recognition and positioning method based on a deep learning visual positioning model provided by the present invention, the preprocessed real-time image is input into the vehicle recognition and positioning model, and a vehicle recognition and positioning result is output, including: Determine the description matching vehicle area existing in the pre-processed real-time image, and output the positioning information, that is, the area with the greatest correlation with the text description; Identify the type of vehicle based on the positioning information and identify the contour information of the vehicle that meets the description in the image; The positioning information, vehicle type and profile information are used as the vehicle identification and positioning result.

[0012] In a second aspect, the present invention further provides a vehicle identification and positioning system based on a deep learning visual positioning model, comprising: A data set construction module, used to collect traffic images, perform data set pre-operation on the traffic images, and construct a training data set; An image acquisition module is used to acquire a real-time image of a vehicle traveling, and preprocess the real-time image of the vehicle traveling to obtain a preprocessed real-time image; A training module, used to construct a multimodal fusion model, and to train and adjust the multimodal fusion model using the image features and text features in the training data set to obtain a vehicle recognition and positioning model; The recognition module is used to input the pre-processed real-time image into the vehicle recognition and positioning model and output the vehicle recognition and positioning result.

[0013] In a third aspect, the present invention also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a vehicle identification and positioning method based on a deep learning visual positioning model as described above is implemented.

[0014] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a vehicle identification and positioning method based on a deep learning visual positioning model as described in any one of the above.

[0015] In a fifth aspect, the present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements a vehicle identification and positioning method based on a deep learning visual positioning model as described above.

[0016] The vehicle identification and positioning method and system based on the deep learning visual positioning model provided by the present invention support deep learning and visual positioning technology by constructing a traffic image dataset, and use a set of keywords to mark the characteristics of each vehicle, thereby realizing the precise association between text description and keywords. This dataset provides strong support for the subsequent vehicle identification and positioning method and system based on the deep learning visual positioning model; fine-tuning the deep learning model to the traffic field not only significantly improves the accuracy of vehicle identification in complex scenarios, but also improves the interpretability of the model, so that the intelligent transportation system can more flexibly respond to various vehicle types and dynamic traffic scenarios. BRIEF DESCRIPTION OF THE DRAWINGS In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 It is a flow chart of a vehicle identification and positioning method based on a deep learning visual positioning model provided by the present invention; Figure 2 This is one of the schematic diagrams of vehicle identification results provided by the present invention; Figure 3 This is the second schematic diagram of the vehicle identification result provided by the present invention; Figure 4 It is a structural schematic diagram of a vehicle identification and positioning system based on a deep learning visual positioning model provided by the present invention; Figure 5 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0019] In view of the defects of the prior art, the present invention proposes a vehicle identification and positioning method based on a deep learning visual positioning model. Figure 1 As shown, including: Step 100: collecting traffic images, performing data set pre-operation on the traffic images, and constructing a training data set; Step 200: Acquire a real-time image of a vehicle traveling, and preprocess the real-time image of the vehicle traveling to obtain a preprocessed real-time image; Step 300: construct a multimodal fusion model, and use the image features and text features in the training data set to train and adjust the multimodal fusion model to obtain a vehicle recognition and positioning model; Step 400: input the pre-processed real-time image into the vehicle recognition and positioning model, and output the vehicle recognition and positioning result.

[0020] In intelligent transportation and monitoring systems, it is often the case that traffic complaint information lacks specific or incorrect license plate information, but contains a large amount of valid vehicle description information, such as vehicle model, color, size, etc. How to make good use of these effective vehicle description information to accurately screen out target vehicles from massive traffic images has become a major challenge at present. In order to solve this problem, the embodiment of the present invention innovatively constructs a traffic image dataset. The dataset contains 1,000 complex traffic images, and the license plates in these images are blurred to protect privacy information while avoiding the use of license plate information. The outline of each vehicle is manually drawn on each image, and a comprehensive set of keywords is defined to describe vehicle features, such as color, vehicle type, relative position information, etc. Based on these keywords, a detailed text description is given to each vehicle, and the features of each vehicle are marked using a set of keywords, realizing the precise association between text description and keywords. This dataset provides strong support for the subsequent vehicle recognition and positioning method and system based on deep learning visual positioning model.

[0021] Specifically, the first step is to build a traffic image dataset for specific scenarios. 1,000 traffic images are collected as the basis of the dataset. The license plates in each traffic image are desensitized or blurred to protect privacy information. The outline of each vehicle is manually drawn in the image for subsequent image processing and recognition.

[0022] Next is the definition of a set of keywords, which are as follows: color (black, white, gray, red, blue, green, yellow, silver, gold), vehicle type (sedan, sport utility vehicle, multi-purpose passenger vehicle, bus, truck, motorcycle, bicycle), location information (intersection, middle of the road), and urban infrastructure (street lights, fire hydrants, traffic lights, traffic signs, public facilities, green belts).

[0023] Provide a text description for each vehicle. Based on the actual situation of each vehicle in the image, use the words in the keyword set to provide a detailed text description for each vehicle. Then organize the processed images and the corresponding text descriptions (keyword set) into a complete data set.

[0024] The second step is to use the constructed scene-specific traffic image dataset to fine-tune the vehicle recognition and positioning method based on the deep learning visual positioning model of the present invention.

[0025] First, camera modules are deployed at traffic monitoring points: camera modules are set up at preset locations on the highway to collect real-time image and video data of vehicles in motion; An image data preprocessing module is established, and a variety of image preprocessing techniques such as denoising, image enhancement and standardization are used to improve image quality and extract basic image features to ensure the stability of subsequent model inputs. In terms of denoising, Gaussian filtering can be used to smooth images and reduce Gaussian noise; median filtering removes "salt and pepper noise" by calculating the median of each pixel area; and bilateral filtering can smooth the interior of the area while maintaining the edge of the image, which is very suitable for processing complex traffic images. In terms of image enhancement, histogram equalization can adjust the contrast of the image and enhance the details; adaptive histogram equalization (AHE) and its improved method CLAHE (contrast-limited adaptive histogram equalization) are suitable for images with uneven lighting, and avoid noise amplification through local equalization or limited contrast enhancement; gamma correction is used to adjust the image brightness to make the details clearer. Finally, in the normalization process, mean-variance normalization normalizes the image pixel values ​​to a distribution with a mean of 0 and a variance of 1 to meet the model input requirements; Min-Max normalization normalizes the pixel values ​​to between 0 and 1 or other specified intervals; color channel normalization normalizes the RGB channels separately to balance the distribution of each channel and reduce the impact of light on the image.

[0026] The image data classification, organization and storage module uses MySQL to establish a traffic image database, which facilitates convenient and efficient management and access to the collected traffic image data.

[0027] The image acquisition module is used to acquire the pre-processed traffic images to be searched with uniform specifications stored in the traffic image database.

[0028] The text acquisition module is used to obtain the text description content of the target vehicle and filter out meaningless character content.

[0029] Establish a text feature encoding module: input natural language description (such as "yellow bus at the intersection" or "black car parked behind the street light"), and convert the description into a high-dimensional semantic feature vector through a pre-trained text encoder. The core of this module is a single-layer bidirectional gated recurrent unit GRU. The gated recurrent unit uses an update gate and a reset gate. The update gate determines how much previous information should be allowed to pass, while the reset gate determines how much previous information should be discarded. The bidirectional gated recurrent unit converts two unidirectional hidden states into a single gate. Each step t is concatenated to form word features. Although a single-layer bidirectional GRU can significantly improve efficiency, when pursuing higher accuracy and more complex text encoding, advanced model architectures such as BERT can also be considered. Through these technical means, the text feature encoding module can accurately capture and transform the semantic information in natural language descriptions, laying a solid foundation for subsequent task processing.

[0030] In the intelligent transportation system, this embodiment fine-tunes the deep learning model for the transportation field, especially in the referring expression segmentation task. Unlike traditional vehicle recognition technology that relies on vehicle detection and classification algorithms, the model of this embodiment is fully adapted to the specific needs of traffic scenes. This includes the ability to handle complex scenes such as mixed vehicles of multiple types, different perspectives and occlusions, which often perform poorly in traditional methods. By combining image features with multimodal information described in natural language, the work of this embodiment significantly improves the accuracy and efficiency of the model in understanding traffic semantics and visual associations. Specifically, this embodiment optimizes the extraction method of semantic features in the deep learning framework, so that it can more accurately capture the subtle meaning of natural language descriptions, and at the same time, combined with specific areas in traffic images, achieve efficient recognition, accurate segmentation and positioning of target vehicles. This deep fine-tuning technology not only significantly improves the accuracy of vehicle recognition in complex scenes, but also improves the interpretability of the model, enabling intelligent transportation systems to more flexibly respond to diverse vehicle types and dynamic traffic scenes.

[0031] Furthermore, a vehicle detection model based on multimodal fusion is constructed. The model is fine-tuned using the constructed scene-specific traffic image dataset, combining image features and text features to accurately match the area in the image corresponding to the description. In terms of the language encoder, in order to demonstrate the effectiveness of SeqTR, this embodiment does not choose to use a pre-trained language encoder (such as BERT), so the language encoder uses a single-layer bidirectional GRU. At each time step t, the bidirectional hidden state is concatenated into , thereby generating word features In terms of the visual encoder, its multi-scale features are down-sampled and flattened from the finest to the coarsest spatial resolution to generate visual features , as input to the fusion module. H (height) and W (width) are 1 / 32 of the original image size respectively. Unlike previous methods, this embodiment only uses the coarsest scale visual features instead of the finest scale features, because this embodiment does not predict binary masks pixel by pixel, which reduces memory consumption during training. In terms of the fusion module, unlike Pix2Seq that only receives pixel input, this embodiment designs a simple and efficient fusion module to align visual and language modalities. Given the visual features and word features First, we construct language features through the maximum pooling of the channel dimension Then, and Perform Hadamard product to generate multimodal features , supply the Transformer encoder: .in, is the tanh function. The word features and visual features are not concatenated and fused using the Transformer encoder, because this will lead to a quadratic increase in complexity. In terms of Transformer and predictor, the standard Transformer encoder updates the multimodal features The feature representation of the target sequence is used, and the decoder predicts the target sequence in an autoregressive manner. The hidden dimension of the Transformer is set to 256, the expansion ratio of the feedforward network is 4, and the number of encoder and decoder layers is 6 and 3 respectively, so the Transformer structure is extremely compact. Since the Transformer is permutation invariant, Sinusoidal positional encoding and learnable positional encoding are added to the input sequence, respectively. To predict the coordinate token, a multi-layer perceptron MLP with a final softmax function is used.

[0032] Finally, a multimodal fusion model is used to detect whether there is a vehicle area matching the description in the input real-time image and output the positioning information; once a vehicle is detected, the model is used to identify the vehicle type (such as SUV, truck, sedan) and mark the specific location of the vehicle that meets the description in the image; the result containing the vehicle location, type and segmentation outline is output in the image or video, which is applied to traffic monitoring, intelligent driving systems and data analysis platforms, such as Figure 2 The identified “yellow bus at the intersection” and Figure 3 The identified “black car parked behind a street light” is shown.

[0033] The vehicle identification and positioning system based on the deep learning visual positioning model provided by the present invention is described below. The vehicle identification and positioning system based on the deep learning visual positioning model described below and the vehicle identification and positioning method based on the deep learning visual positioning model described above can be referenced to each other.

[0034] Figure 4 is a schematic diagram of the structure of a vehicle identification and positioning system based on a deep learning visual positioning model provided by an embodiment of the present invention, such as Figure 4 As shown, it includes: a data set construction module 41, an image acquisition module 42, a training module 43 and a recognition module 44, wherein: The data set construction module 41 is used to collect traffic images, perform data set pre-operation on the traffic images, and construct a training data set; the image acquisition module 42 is used to obtain real-time images of vehicle driving, pre-process the real-time images of vehicle driving, and obtain pre-processed real-time images; the training module 43 is used to construct a multimodal fusion model, and use the image features and text features in the training data set to train and adjust the multimodal fusion model to obtain a vehicle recognition and positioning model; the recognition module 44 is used to input the pre-processed real-time image into the vehicle recognition and positioning model, and output the vehicle recognition and positioning result.

[0035] Figure 5 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 5 As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530 and a communication bus 540, wherein the processor 510, the communication interface 520 and the memory 530 communicate with each other through the communication bus 540. The processor 510 may call the logic instructions in the memory 530 to execute a vehicle identification and positioning method based on a deep learning visual positioning model, the method comprising: collecting traffic images, performing data set pre-operation on the traffic images, and constructing a training data set; acquiring real-time images of vehicle driving, pre-processing the real-time images of vehicle driving, and obtaining pre-processed real-time images; constructing a multimodal fusion model, using the image features and text features in the training data set to train and adjust the multimodal fusion model, and obtaining a vehicle identification and positioning model; inputting the pre-processed real-time image into the vehicle identification and positioning model, and outputting a vehicle identification and positioning result.

[0036] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0037] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the vehicle identification and positioning method based on the deep learning visual positioning model provided by the above methods. The method includes: collecting traffic images, performing data set pre-operation on the traffic images, and constructing a training data set; obtaining real-time images of vehicle driving, pre-processing the real-time images of vehicle driving, and obtaining pre-processed real-time images; constructing a multimodal fusion model, using the image features and text features in the training data set to train and adjust the multimodal fusion model to obtain a vehicle identification and positioning model; inputting the pre-processed real-time image into the vehicle identification and positioning model, and outputting the vehicle identification and positioning result.

[0038] On the other hand, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by the processor, it is implemented to execute the vehicle identification and positioning method based on the deep learning visual positioning model provided by the above methods, the method comprising: collecting traffic images, performing data set pre-operation on the traffic images, and constructing a training data set; obtaining real-time images of vehicle driving, pre-processing the real-time images of vehicle driving, and obtaining pre-processed real-time images; constructing a multimodal fusion model, using the image features and text features in the training data set to train and adjust the multimodal fusion model, and obtain a vehicle identification and positioning model; inputting the pre-processed real-time image into the vehicle identification and positioning model, and outputting a vehicle identification and positioning result. The device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. Ordinary technicians in this field can understand and implement without creative labor.

[0039] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0040] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A vehicle identification and positioning method based on a deep learning visual positioning model, characterized in that: include: Collecting traffic images, performing data set pre-operation on the traffic images, and constructing a training data set; Acquire a real-time image of a vehicle traveling, and preprocess the real-time image of the vehicle traveling to obtain a preprocessed real-time image; Constructing a multimodal fusion model, and using the image features and text features in the training data set to train and adjust the multimodal fusion model to obtain a vehicle recognition and positioning model; The pre-processed real-time image is input into the vehicle recognition and positioning model, and a vehicle recognition and positioning result is output.

2. The vehicle identification and positioning method based on the deep learning visual positioning model according to claim 1 is characterized in that: Collecting traffic images, performing data set pre-operation on the traffic images, and constructing a training data set, including: Collect several traffic images; Performing fuzzy processing on the license plate in each traffic image to obtain a fuzzy processed image; Manually draw the outline of each vehicle on each blurred image to obtain a vehicle outline image; defining a keyword set, the keyword set comprising color, vehicle type, location information, and urban infrastructure; Writing a text description for each vehicle based on the words in the keyword set; Combine the processed images with the text description of each vehicle to construct a training dataset.

3. The vehicle identification and positioning method based on the deep learning visual positioning model according to claim 1 is characterized in that: Acquiring a real-time image of a vehicle driving, and preprocessing the real-time image of the vehicle driving to obtain a preprocessed real-time image, including: Deploy camera modules at preset traffic monitoring points to collect real-time image and video data of vehicles in motion; Gaussian filtering is used to remove Gaussian noise in the image, median filtering is used to calculate the median of each pixel area to remove salt and pepper noise in the image, and bilateral filtering is used to maintain image edge information and smooth the interior of the image area to obtain a denoised image; Adaptive histogram equalization and improved adaptive histogram equalization are used to limit the contrast enhancement of the image, and gamma correction is used to adjust the image brightness to obtain the enhanced image; Mean-variance normalization is used to normalize the image pixel values ​​to a distribution with a mean of 0 and a variance of 1. Min-Max normalization is used to normalize the image pixel values ​​to between 0 and 1 or other specified intervals. Color channel normalization is used to normalize the RGB channels of the image separately to balance the distribution of each channel to obtain a standardized image.

4. The vehicle identification and positioning method based on the deep learning visual positioning model according to claim 1 is characterized in that: Acquiring a real-time image of a vehicle driving, preprocessing the real-time image of the vehicle driving, and obtaining the preprocessed real-time image, further comprising: Use MySQL to build a traffic image database; An image acquisition module is established to acquire the pre-processed real-time image.

5. The vehicle identification and positioning method based on the deep learning visual positioning model according to claim 1 is characterized in that: Constructing a multimodal fusion model, using the image features and text features in the training data set to train and adjust the multimodal fusion model, and obtaining a vehicle recognition and positioning model, including: Establish a text feature encoding module, the text feature encoding module includes a single-layer bidirectional gated recurrent unit, input the natural language description into the text feature encoding module, convert it into a high-dimensional semantic feature vector, use the text feature encoding module as a language encoder, splice the bidirectional hidden states in each time step, and generate word features; In the visual encoder, multi-scale features are down-sampled and flattened from the finest to the coarsest spatial resolution to generate visual features; The word features and the visual features are input into a fusion module, language features are constructed through maximum pooling of the channel dimension, Hadamard products are performed on the word features and the visual features to generate multimodal features, and the multimodal features are input into a Transformer encoder to obtain the vehicle recognition and positioning model.

6. The vehicle identification and positioning method based on the deep learning visual positioning model according to claim 1 is characterized in that: Inputting the preprocessed real-time image into the vehicle recognition and positioning model, and outputting the vehicle recognition and positioning result, including: Determine the description matching vehicle area existing in the pre-processed real-time image, and output positioning information; Identify the type of vehicle based on the positioning information and identify the contour information of the vehicle that meets the description in the image; The positioning information, vehicle type and profile information are used as the vehicle identification and positioning result.

7. A vehicle identification and positioning system based on a deep learning visual positioning model, characterized in that: include: A data set construction module, used to collect traffic images, perform data set pre-operation on the traffic images, and construct a training data set; An image acquisition module is used to acquire a real-time image of a vehicle traveling, and preprocess the real-time image of the vehicle traveling to obtain a preprocessed real-time image; A training module, used to construct a multimodal fusion model, and to train and adjust the multimodal fusion model using the image features and text features in the training data set to obtain a vehicle recognition and positioning model; The recognition module is used to input the pre-processed real-time image into the vehicle recognition and positioning model and output the vehicle recognition and positioning result.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the vehicle identification and positioning method based on the deep learning visual positioning model as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the vehicle identification and positioning method based on the deep learning visual positioning model as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the vehicle identification and positioning method based on the deep learning visual positioning model as described in any one of claims 1 to 6 is implemented.