Cross-modal retrieval method for traffic pictures
By constructing a cross-modal retrieval method, converting images into vector data and outputting the results in a unified vector space, the problem of difficult alignment of image and text data is solved, and efficient traffic image retrieval and management is achieved.
Patent Information
- Application Number
- CN202510996781.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-09-12
AI Technical Summary
When processing image and text data, existing technologies find it difficult to effectively align the semantic information of the two, resulting in low cross-modal retrieval accuracy, high false detection rate, and easy omission or misjudgment.
A cross-modal retrieval method for traffic images is constructed. The first model converts images into image vectors, and the text encoder and image encoder are combined to output the results in a unified vector space. The model is optimized using a multi-task classification layer, a cross-entropy loss function, and the AdamW optimizer to reduce the number of parameters and prevent overfitting. The cosine distance is used to retrieve matching image indexes.
It can quickly locate the trajectory or violation record of the target vehicle, provide strong evidence, enhance the deterrent effect of traffic laws and regulations, reduce manual intervention, lower the error rate, and improve the level of intelligent traffic management.
Smart Images

Figure CN120632133A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of big data and relates to image retrieval technology, specifically a cross-modal retrieval method for traffic images. Background Art
[0002] With the widespread use of the internet, social media, and digital devices, multimodal data is experiencing explosive growth. In the transportation sector, especially, highway surveillance systems generate billions of high-definition images and videos daily. Combining this data with the text of traffic violation regulations could theoretically enable automated and intelligent violation identification and evidence extraction through cross-modal retrieval technology. However, existing technologies face significant bottlenecks in processing this type of multimodal data.
[0003] In existing technologies, when processing image and text data, since image and text data belong to different modalities and their underlying feature spaces are very different, traditional single-modal analysis methods (such as pure image recognition or text matching) find it difficult to effectively align the semantic information of the two, resulting in low cross-modal retrieval accuracy, high false detection rate, and easy occurrence of technical problems such as missed or misjudgment.
[0004] The present invention provides a cross-modal retrieval method for traffic images to solve the above technical problems. Summary of the Invention
[0005] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, the present invention proposes a cross-modal retrieval method for traffic images, which is used to solve the technical problem that the single-modal analysis method in the prior art is difficult to effectively align the semantic information of the two, resulting in low cross-modal retrieval accuracy, high false detection rate, and easy omission or misjudgment.
[0006] To achieve the above objectives, a first aspect of the present invention provides a cross-modal retrieval method for traffic images, comprising: Acquire a first image of a vehicle to be found and vehicle information of the vehicle to be found; Inputting the first image into a first model to obtain a first image vector of the first image; wherein the first model is a model trained based on images of multiple vehicles and image vectors corresponding to the images of the multiple vehicles; Retrieving a first image index that matches both the first image vector and the vehicle information from a traffic image database; wherein the traffic image database includes: a plurality of traffic images and an image index corresponding to each traffic image, wherein the image index corresponding to each traffic image is an image index generated based on an image vector of the traffic image and the vehicle information in the traffic image; and wherein the image vector of the traffic image is an image vector obtained by inputting the first model; The traffic image corresponding to the first image index is determined as the traffic image of the vehicle to be found.
[0007] Preferably, the traffic image database is generated based on the following process: Inputting a plurality of traffic images with vehicle positions pre-marked into a first model to obtain an image vector of each vehicle in the plurality of traffic images; Inputting the plurality of traffic images into a second model to obtain vehicle information of each vehicle in the plurality of traffic images; wherein the second model is a model trained based on the plurality of vehicle images and the vehicle information; constructing an image index for each traffic image based on the image vector and vehicle information of each vehicle in each traffic image; Based on each traffic image and the image index of each traffic image, the traffic image database.
[0008] Preferably, the method for acquiring the plurality of traffic images with the vehicle positions marked includes: Acquire multiple initial traffic images, and input the multiple initial traffic images into a third model to obtain the multiple traffic images; wherein the third model is a model for marking vehicle positions in images obtained by training based on multiple historical traffic images.
[0009] Preferably, the training process of the second model includes: Retrieving multiple traffic images, scaling their pixel values to a preset fixed range, unifying the input size of the traffic images, and performing image enhancement on the multiple traffic images. Image enhancement includes geometric transformation, brightness control within a set range, and contrast adjustment. The backbone network is constructed based on a pre-trained residual neural network. A lightweight attention mechanism is embedded in the residual blocks of the residual neural network to enhance the feature response of color-sensitive areas to obtain an amplified color channel. A multi-task classification layer is constructed based on the color channel to obtain vehicle information. The vehicle information includes vehicle color, vehicle classification, and vehicle orientation. The steps of embedding a lightweight attention mechanism in the residual block of the residual neural network to enhance the feature response of the color-sensitive area to obtain the amplified color channel are as follows: compressing the spatial dimension through global average pooling to generate a channel description vector; learning the relationship between the channel description vectors through a fully connected layer to generate channel weights; and multiplying the channel weights by the feature map of the traffic image channel by channel to obtain the amplified color channel.
[0010] Preferably, constructing a multi-task classification layer based on the color channel to obtain vehicle information includes: The calculation formula for the multi-task classification layer is constructed: Through the equation Calculate a single vector consisting of multi-task prediction scores; the multi-task includes: color, vehicle type, and orientation; x is the global average pooling output of the last layer of the backbone network; represents the color bias, represents the bias term of the vehicle model, represents the bias term of the orientation, represents the weight matrix of the color, represents the weight matrix of the vehicle model, The weight matrix representing the orientation; is the prediction score of color, is the prediction score of the vehicle model, is the predicted score for the orientation; The updated parameters are obtained by combining the cross entropy loss function with the AdamW optimizer function. The parameters include weight decay strength and adaptive gradient terms. Preset initial learning rate , through the formula Calculate the learning rate when the number of iterations is t; where, is the minimum learning rate, is the maximum learning rate, is the maximum number of iterations; the value range of t is a positive integer; if , then the iteration process is stopped, and the maximum value of the multi-task prediction score in the iteration process is used as the corresponding vehicle information.
[0011] The present invention constructs a multi-task classification layer to adapt to multi-task output requirements (color category, vehicle model, and orientation classification), while reducing the number of model parameters and improving deployment efficiency. It uses the cross-entropy loss function and AdamW optimizer to prevent overfitting of prediction scores.
[0012] Preferably, the cross entropy loss function is obtained by: The cross entropy loss function is constructed as: Where: The unique hot encoding preset for the true label of the i-th category, is the probability of the i-th category predicted by the model; = , = , = .
[0013] Preferably, the method for obtaining the AdamW optimizer function includes: Construct the AdamW optimizer function as follows: ;in, is the updated parameter when the number of iterations is t+1, is the parameter corresponding to the current iteration number t, is the first-order momentum after the preset bias correction, is the bias-corrected second-order momentum, is the preset stable item, Weight decay coefficient.
[0014] Need to explain, the preset stability items The value range is (0, ), so that the denominator of the AdamW optimizer function is non-zero.
[0015] Preferably, the updated parameters obtained by combining the cross entropy loss function with the AdamW optimizer function include: Combine the cross entropy loss function with the AdamW optimizer function: ;in, is the preset weight attenuation strength, is the adaptive gradient term of the loss with respect to the model parameters.
[0016] Preferably, the method for obtaining the first model includes: The images of multiple vehicles are respectively subjected to text encoding and image encoder to obtain vector representation of the text and vector representation of images ,in, is the Nth text encoding value, For the Nth image encoding value, the vector representation of the text and the vector of the picture are combined into an N-dimensional matrix; The first model is constructed based on the N-dimensional matrix: Where, Represents the image vector of the Nth vehicle to be found.
[0017] The present invention converts images of multiple vehicles into vector data by constructing a first model, and converts unstructured image heterogeneous data into measurable vector data.
[0018] Preferably, retrieving a first image index that matches both the first image vector and the vehicle information from a traffic image database includes: Retrieving the first image vector and the n-dimensional matrix of the vector representation of the image of the vehicle to be searched; By formula The cosine distance value is calculated; where, is the vector value of the i-th dimension in the first image, is the vector value of the i-th dimension in the image of the vehicle to be found in the traffic image database; An image vector with the smallest cosine distance value is retrieved, and the image vector of the traffic image is matched with vehicle information in the traffic image to generate an image index as a first image index.
[0019] Compared with the prior art, the present invention has the following beneficial effects: 1. The present invention can retrieve target vehicle traffic pictures by uploading the picture of the target vehicle to find all target traffic pictures, quickly locate the trajectory or violation record of the target vehicle and other information, provide strong evidence for traffic law enforcement, enhance the deterrent effect of traffic laws and regulations, promote the standardization of traffic order, reduce manual intervention, reduce labor costs and error rates, and improve the overall intelligent level of traffic management, so that traffic management departments can more efficiently respond to increasingly complex traffic conditions and management tasks.
[0020] 2. This paper constructs a first model that combines a text encoder with an image encoder, outputting results in a unified vector space. The distance between the text vector and the image vector reflects the closeness of their semantic association. The text passes through the text transformer, and the image encoder uses CNN's ResNet. By learning from a large number of image and text pairs, the system can understand and classify new images even without targeted task training. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0022] Figure 1 Schematic diagram of the cross-modal search process of the present invention; Figure 2 A schematic diagram of the specific steps of vehicle image information recognition according to the present invention; Figure 3 Schematic diagram of the specific steps of vehicle image vectorization of the present invention; DETAILED DESCRIPTION
[0023] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0024] See also Figure 1The first embodiment of the present invention provides a cross-modal retrieval method for traffic images, comprising: Acquire a first image of a vehicle to be found and vehicle information of the vehicle to be found; Inputting the first image into a first model to obtain a first image vector of the first image; wherein the first model is a model trained based on images of multiple vehicles and image vectors corresponding to the images of the multiple vehicles; Retrieving a first image index that matches both the first image vector and the vehicle information from a traffic image database; wherein the traffic image database includes: a plurality of traffic images and an image index corresponding to each traffic image, wherein the image index corresponding to each traffic image is an image index generated based on an image vector of the traffic image and the vehicle information in the traffic image; and wherein the image vector of the traffic image is an image vector obtained by inputting the first model; The traffic image corresponding to the first image index is determined as the traffic image of the vehicle to be found.
[0025] See also Figure 2 The specific steps of vehicle image information recognition are as follows: obtaining multiple initial traffic images, inputting the multiple initial traffic images into a third model to obtain multiple traffic images; wherein the third model is a model for marking vehicle positions in an image obtained by training based on multiple historical traffic images; Retrieving multiple traffic images, scaling their pixel values to a preset fixed range, unifying the input size of the traffic images, and performing image enhancement on the multiple traffic images. Image enhancement includes geometric transformation, brightness control within a set range, and contrast adjustment. The backbone network is constructed based on a pre-trained residual neural network. A lightweight attention mechanism is embedded in the residual blocks of the residual neural network to enhance the feature response of color-sensitive areas to obtain an amplified color channel. A multi-task classification layer is constructed based on the color channel to obtain vehicle information. The vehicle information includes vehicle color, vehicle classification, and vehicle orientation. For example: training data is collected from vehicle images taken at highway toll booths; Image preprocessing and enhancement: First, scale the pixel values to [0, 1] or [-1, 1], and unify the image input size to 256 256, random cropping, rotation, and flipping of sample data, and color perturbation with brightness within the range [0.8, 1.2] and contrast adjustment [0.8, 1.2] (to avoid affecting color labels). Finally, RGB can be experimentally converted to HSV / HSL, and the hue channel can be used to enhance color features or as an additional input channel; The backbone network is based on the pre-trained residual neural network ResNet34, which has powerful feature extraction capabilities and moderate computational complexity (balancing performance and computational complexity). It also introduces an efficient attention mechanism and embeds a lightweight attention module in the ResNet34 residual block to enhance the feature response to color-sensitive areas.
[0026] Optimize and solve the problem of vanishing gradients in deep networks: , thereby improving the network's attention to the color, model and orientation area and suppressing interference.
[0027] Create a multi-task classification layer: in, ; color: color, type: vehicle type, orientation: direction; Color dimension (red, blue, green, etc.), : Vehicle type dimension (sedan, SUV, etc.), : Toward dimension (ahead, backward); x: Global average pooling (GAP) output from the last layer of the ResNet34 backbone network, W: weight matrix of color, type, and orientation, b: bias term of color, type, and orientation; y: Joint output: The predicted scores of color, vehicle type, and orientation are spliced into a single vector to facilitate multi-task joint optimization. , Model: , towards: .
[0028] The loss function uses a joint loss, and the cross entropy loss is calculated for color, vehicle type, and orientation respectively and then weighted summed: L=α Lcolor +β Ltype +γ Lorientation , α, β, γ are adjusted according to the importance of the task (e.g. α=1.0, β=0.8, γ=0.5) The cross-entropy loss function and AdamW optimizer are combined with weight decay to prevent overfitting, adapt to multi-task output requirements (color category, vehicle model, and orientation classification), reduce the number of model parameters, and improve deployment efficiency. The cross entropy loss function is: ; yi: the unique hot encoding of the true label (i-th class is 1, the rest are 0), pi: the probability of the i-th class predicted by the model; By minimizing the cross entropy loss, the model learns a probability distribution that is consistent with the true label. Although cross entropy itself does not directly regularize the model, the optimization process combined with weight decay and AdamW's update strategy can avoid the model's overconfidence in the training data (such as the predicted probability is too sharp). AdamW optimizer: ;in, is the updated parameter, Current parameters, Learning rate 0.001, Bias-corrected first-order momentum, Bias-corrected second-order momentum, The minimum constant of the numerical stability term (usually set to 1×10^ 8) Prevent the denominator from being zero, Weight decay coefficient 0.01 Combining the fork entropy loss function with the AdamW optimizer: ; The weight decay strength is set to 0.01, The gradient of loss to model parameters, AdamW is based on the gradient Calculate adaptive learning rate and momentum terms; Dynamically adjust the learning rate to prevent local optimality and improve the convergence speed; through the formula Calculate the learning rate when the number of iterations is t; where, is the minimum learning rate, is the maximum learning rate, is the maximum number of iterations; Monitor the validation set loss: terminate if the number of iterations does not decrease for three consecutive times; After the above steps, a vehicle information recognition model is trained. Based on the image obtained from the vehicle position marked in the above steps, the vehicle information including vehicle color, vehicle type, vehicle direction and other information is recognized.
[0029] See also Figure 3 The specific steps of vehicle image vectorization are: multiple vehicle images are respectively subjected to text encoding and image encoder to obtain vector representation of text and vector representation of images ,in, is the Nth text encoding value, For the Nth image encoding value, the vector representation of the text and the vector of the picture are combined into an N-dimensional matrix; The first model is constructed based on the N-dimensional matrix: Where, Represents the image vector of the Nth vehicle to be found; An n-dimensional matrix represented by the first image vector and the vector of the image of the vehicle to be found; By formula The cosine distance value is calculated; where, is the vector value of the i-th dimension in the first image, is the vector value of the i-th dimension in the image of the vehicle to be found in the traffic image database; An image vector with the smallest cosine distance value is retrieved, and the image vector of the traffic image is matched with vehicle information in the traffic image to generate an image index as a first image index; and the traffic image corresponding to the first image index is determined as the traffic image of the vehicle to be found.
[0030] For example, we build an image vectorization model that combines a text encoder with an image encoder, outputting the result in a unified vector space. The distance between the text vector and the image vector reflects the closeness of their semantic connection. The text passes through the text transformer, and the image encoder uses a CNN ResNet. By learning from a large number of image and text pairs, it can understand and classify new images even without specific task training. Training process: A batch of data sets has N samples, so for this batch, after text encoding and image encoder, we can get and , then this can get N The matrix of N is similar to the attention map; it is calculated using cosine distance. Then we get N^2 results, of which N are positive samples, so we hope that their cosine distance is 1, and the remaining N^2 For N negative sample pairs, we hope that their cosine distance is -1, so the optimization goal is set as: ; Thus, an image vectorization model is obtained, and the vehicle image information is converted into vector data through the image vectorization model. After this step, the unstructured image heterogeneous data is converted into measurable vector data.
[0031] Image vector storage: Integrate acquired vehicle information data and obtained vehicle vector data, create a BinaryIVF vector index, and save it in the vector database. This completes the conversion of unstructured data into structured data. BinaryIVF is an efficient vector retrieval technology that combines binary vectors and inverted file indexes (IVF). Image retrieval: Based on the vehicle image, the vectorized information of the vehicle image is obtained and the cosine similarity is calculated with the vectors that meet the conditions in the database: ; Retrieve the top N images with the highest similarity (i.e., the images corresponding to the minimum cosine similarity) to find the target vehicle image.
[0032] To sum up, the retrieval of target vehicle traffic pictures can find all target traffic pictures by uploading pictures of the target vehicle, quickly locate the trajectory or violation record of the target vehicle and other information, provide strong evidence for traffic law enforcement, enhance the deterrent effect of traffic laws and regulations, promote the standardization of traffic order, reduce manual intervention, reduce labor costs and error rates, and improve the overall intelligent level of traffic management, so that traffic management departments can respond to increasingly complex traffic conditions and management tasks more efficiently.
[0033] Some of the data in the above formula are calculated by removing the dimensions and taking their numerical values. The formula is a formula that is closest to the actual situation obtained by software simulation of a large amount of collected data; the preset parameters and preset thresholds in the formula are set by technical personnel in this field according to actual conditions or obtained through simulation of a large amount of data.
[0034] The working principle of the present invention is as follows: the present invention obtains a first image of a vehicle to be found and vehicle information of the vehicle to be found; inputs the first image into a first model to obtain a first image vector of the first image; retrieves a first image index that matches both the first image vector and the vehicle information from a traffic image database; and determines the traffic image corresponding to the first image index as the traffic image of the vehicle to be found.
[0035] The above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.
Claims
1. A cross-modal retrieval method for traffic images, characterized in that: include: Acquire a first image of a vehicle to be found and vehicle information of the vehicle to be found; Inputting the first image into a first model to obtain a first image vector of the first image; wherein the first model is a model trained based on images of multiple vehicles and image vectors corresponding to the images of the multiple vehicles; Retrieving a first image index that matches both the first image vector and the vehicle information from a traffic image database; wherein the traffic image database includes: a plurality of traffic images and an image index corresponding to each traffic image, wherein the image index corresponding to each traffic image is an image index generated based on an image vector of the traffic image and the vehicle information in the traffic image; and wherein the image vector of the traffic image is an image vector obtained by inputting the first model; The traffic image corresponding to the first image index is determined as the traffic image of the vehicle to be found.
2. A cross-modal retrieval method for traffic images according to claim 1, characterized in that: The traffic image database is generated based on the following process: Inputting a plurality of traffic images with vehicle positions pre-marked into a first model to obtain an image vector of each vehicle in the plurality of traffic images; Inputting the plurality of traffic images into a second model to obtain vehicle information of each vehicle in the plurality of traffic images; wherein the second model is a model trained based on the plurality of vehicle images and the vehicle information; constructing an image index for each traffic image based on the image vector and vehicle information of each vehicle in each traffic image; Based on each traffic image and the image index of each traffic image, the traffic image database.
3. A cross-modal retrieval method for traffic images according to claim 2, characterized in that: The method for acquiring the plurality of traffic images with the vehicle positions marked includes: Acquire multiple initial traffic images, and input the multiple initial traffic images into a third model to obtain the multiple traffic images; wherein the third model is a model for marking vehicle positions in images obtained by training based on multiple historical traffic images.
4. A cross-modal retrieval method for traffic images according to claim 2, characterized in that: The training process of the second model includes: Retrieving multiple traffic images, scaling their pixel values to a preset fixed range, unifying the input size of the traffic images, and performing image enhancement on the multiple traffic images. Image enhancement includes geometric transformation, brightness control within a set range, and contrast adjustment. The backbone network is constructed based on a pre-trained residual neural network, and a lightweight attention mechanism is embedded in the residual block of the residual neural network to enhance the feature response of color-sensitive areas to obtain the amplified color channel; a multi-task classification layer is constructed based on the color channel to obtain vehicle information; the vehicle information includes: vehicle color, vehicle classification and vehicle orientation.
5. A cross-modal retrieval method for traffic images according to claim 4, characterized in that: The multi-task classification layer is constructed based on the color channel to obtain vehicle information, including: The calculation formula for the multi-task classification layer is constructed: Through the equation Calculate a single vector consisting of multi-task prediction scores; the multi-task includes: color, vehicle type, and orientation; x is the global average pooling output of the last layer of the backbone network; represents the color bias, represents the bias term of the vehicle model, represents the bias term of the orientation, represents the weight matrix of the color, represents the weight matrix of the vehicle model, The weight matrix representing the orientation; is the prediction score of color, is the prediction score of the vehicle model, is the predicted score for the orientation; The updated parameters are obtained by combining the cross entropy loss function with the AdamW optimizer function. The parameters include weight decay strength and adaptive gradient terms. Preset initial learning rate , through the formula Calculate the learning rate when the number of iterations is t; where, is the minimum learning rate, is the maximum learning rate, is the maximum number of iterations; the value range of t is a positive integer; if , then the iteration process is stopped, and the maximum value of the multi-task prediction score in the iteration process is used as the corresponding vehicle information.
6. A cross-modal retrieval method for traffic images according to claim 5, characterized in that: The method of obtaining the cross entropy loss function includes: The cross entropy loss function is constructed as: Where: The unique hot encoding preset for the true label of the i-th category, is the probability of the i-th category predicted by the model; = , = , = .
7. A cross-modal retrieval method for traffic images according to claim 5, characterized in that: The method for obtaining the AdamW optimizer function includes: Construct the AdamW optimizer function as follows: ;in, is the updated parameter when the number of iterations is t+1, is the parameter corresponding to the current iteration number t, is the first-order momentum after the preset bias correction, is the bias-corrected second-order momentum, is the preset stable item, Weight decay coefficient.
8. A cross-modal retrieval method for traffic images according to claim 5, characterized in that: The updated parameters obtained by combining the cross entropy loss function with the AdamW optimizer function include: Combine the cross entropy loss function with the AdamW optimizer function: ;in, is the preset weight attenuation strength, is the adaptive gradient term of the loss with respect to the model parameters.
9. A cross-modal retrieval method for traffic images according to claim 1, characterized in that: The method for obtaining the first model includes: The images of multiple vehicles are respectively subjected to text encoding and image encoder to obtain vector representation of the text and vector representation of images ,in, is the Nth text encoding value, For the Nth image encoding value, the vector representation of the text and the vector of the picture are combined into an N-dimensional matrix; The first model is constructed based on the N-dimensional matrix: Where, Represents the image vector of the Nth vehicle to be found.
10. A cross-modal retrieval method for traffic images according to claim 1, characterized in that: The step of retrieving a first image index that matches both the first image vector and the vehicle information from a traffic image database includes: Retrieving the first image vector and the n-dimensional matrix of the vector representation of the image of the vehicle to be searched; By formula The cosine distance value is calculated; where, is the vector value of the i-th dimension in the first image, is the vector value of the i-th dimension in the image of the vehicle to be found in the traffic image database; An image vector with the smallest cosine distance value is retrieved, and the image vector of the traffic image is matched with vehicle information in the traffic image to generate an image index as a first image index.