Vehicle query method and vehicle query device
By combining reinforcement learning models and lightweight detection models, efficient vehicle querying of multi-camera monitoring networks was achieved, solving the problem of high data processing computational overhead and improving query accuracy and efficiency.
Patent Information
- Application Number
- CN202510811152.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-11-18
AI Technical Summary
The data processing and computational overhead of multi-camera surveillance networks is high, and existing technologies have not been able to effectively solve this problem.
A reinforcement learning model is used to determine the filter combination, and vehicle video data is filtered in multiple dimensions. A feature library is built and matched by combining a lightweight object detection model and a vector retrieval index structure, thereby reducing the amount of irrelevant data processing.
It reduces the computational overhead of multi-camera monitoring networks, improves query accuracy and efficiency, and adapts to real-time or near-real-time query needs.
Smart Images

Figure CN120973828A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of vehicle detection, and in particular to a vehicle query method and a vehicle query device. BACKGROUND
[0002] An intelligent transportation system (ITS) is an effective means for improving traffic efficiency, ensuring driving safety, and reducing traffic congestion in a traffic management scenario. Vehicle tracking technology is one of the key technologies involved in an intelligent transportation system. In order to overcome the problem that a traditional single camera has a limited tracking perspective and is difficult to achieve comprehensive tracking of vehicles in a complex traffic environment, in related technologies, multiple cameras deployed at different positions are used to form a monitoring network covering a wide area, so as to achieve continuous tracking and monitoring of vehicles.
[0003] Since a multi-camera monitoring network needs to process a large amount of image or video data in real time or quasi-real time, the computational overhead of multi-camera data is large.
[0004] At present, there is no effective solution to the problem of large data processing computational overhead of a multi-camera monitoring network in related technologies. SUMMARY
[0005] A vehicle query method and a vehicle query device are provided in the present embodiment to solve the problem of large data processing computational overhead of a multi-camera monitoring network in related technologies.
[0006] In a first aspect, a vehicle query method is provided in the present embodiment, comprising:
[0007] obtaining a target vehicle image to be queried;
[0008] determining a filter combination for vehicle video data according to a trained reinforcement learning model; the filter combination comprises filter conditions for different dimensions;
[0009] filtering the vehicle video data based on the filter combination and at least according to the target vehicle image, to obtain filtered video data related to the target vehicle image;
[0010] performing vehicle detection on the filtered video data, extracting vehicle features of detected vehicle targets, and constructing a feature library according to the vehicle features;
[0011] matching target vehicle features in the target vehicle image with each vehicle feature in the feature library, and determining a candidate vehicle list matched with the target vehicle image according to a matching result.
[0012] In some embodiments, determining a filter combination for vehicle video data according to a trained reinforcement learning model comprises:
[0013] Based on the trained reinforcement learning model, and according to the changes in the video characteristics of the vehicle video data, a corresponding filter combination is determined for each video segment in the vehicle video data.
[0014] In some embodiments, based on the filter combination, the vehicle video data is filtered at least according to the target vehicle image to obtain filtered video data related to the target vehicle image, including:
[0015] Based on the filter combination, the vehicle video data is filtered according to the target vehicle image and the set query conditions to obtain filtered video data related to the target vehicle image.
[0016] In some embodiments, vehicle detection is performed on the filtered video data, vehicle features are extracted from the detected vehicle targets, and a feature library is constructed based on the vehicle features, including:
[0017] Vehicle targets are obtained by detecting vehicles in the filtered video data according to a preset target detection model.
[0018] Based on a preset target extraction model, vehicle features are extracted from the vehicle target.
[0019] In some embodiments, vehicle detection is performed on the filtered video data according to a preset target detection model to obtain vehicle targets, including:
[0020] Based on the target detector interface, vehicle targets are obtained by using a pre-selected target detection model to detect vehicles in the filtered video data.
[0021] In some embodiments, the method further includes:
[0022] The vehicle targets detected in the filtered video data are tracked to obtain tracking information.
[0023] In some embodiments, after obtaining the candidate vehicle list, the method further includes:
[0024] Based on the tracking information, the spatiotemporal information of each vehicle target in the candidate vehicle list is determined.
[0025] In some embodiments, vehicle detection is performed on the filtered video data, vehicle features are extracted from the detected vehicle targets, and a feature library is constructed based on the vehicle features, including:
[0026] All vehicle features in the feature library are stored in a preset vector retrieval index structure.
[0027] In some embodiments, the target vehicle features in the target vehicle image are matched with vehicle features in the feature library, and a list of candidate vehicles matching the target vehicle image is determined based on the matching results, including:
[0028] The query method of the vector retrieval index structure is invoked to match the stored vehicle features according to the target vehicle features, and the matching results are obtained.
[0029] Secondly, this embodiment provides a vehicle query device, including: an acquisition module, a determination module, a filtering module, a detection and extraction module, and a matching module; wherein:
[0030] The acquisition module is used to acquire the image of the target vehicle to be queried;
[0031] The determining module is used to determine the filter combination for vehicle video data based on the trained reinforcement learning model.
[0032] The filtering module is used to filter the vehicle video data based on the filter combination, at least according to the target vehicle image, to obtain filtered video data related to the target vehicle image;
[0033] The detection and extraction module is used to detect vehicles in the filtered video data, extract vehicle features from the detected vehicle targets, and construct a feature library based on the vehicle features.
[0034] The matching module is used to match the target vehicle features in the target vehicle image with the vehicle features in the feature library, and determine a list of candidate vehicles that match the target vehicle image based on the matching results.
[0035] Compared with related technologies, this embodiment provides a vehicle query method and a vehicle query device. The vehicle query method involves: acquiring an image of the target vehicle to be queried; determining a filter combination for vehicle video data based on a trained reinforcement learning model; the filter combination including filtering conditions for different dimensions; filtering the vehicle video data based on the filter combination, at least according to the target vehicle image, to obtain filtered video data related to the target vehicle image; performing vehicle detection on the filtered video data, extracting vehicle features from the detected vehicle targets, and constructing a feature library based on the vehicle features; matching the target vehicle features in the target vehicle image with the vehicle features in the feature library, and determining a list of candidate vehicles matching the target vehicle image based on the matching results. This method filters vehicle videos collected by multiple camera devices, thereby reducing the amount of data required for subsequent target detection processing and reducing computational overhead.
[0036] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description
[0037] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0038] Figure 1 This is a hardware structure block diagram of the terminal for the vehicle query method according to an embodiment of this application;
[0039] Figure 2 This is a flowchart of the vehicle query method according to an embodiment of this application;
[0040] Figure 3 This is a schematic diagram illustrating the training and application of a reinforcement learning agent according to an embodiment of this application;
[0041] Figure 4 This is a flowchart of a vehicle query method according to some embodiments of this application;
[0042] Figure 5 This is a structural block diagram of the vehicle query device according to an embodiment of this application. Detailed Implementation
[0043] To better understand the purpose, technical solution, and advantages of this application, the application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0044] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning understood by one of ordinary skill in the art to which this application pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these” used in this application do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to these processes, methods, products, or devices. Words such as “connected,” “linked,” and “coupled” used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. Normally, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," "third," etc., used in this application are merely to distinguish similar objects and do not represent a specific order of objects.
[0045] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. For example, it can run on a terminal. Figure 1 This is a hardware structure block diagram of the terminal for the vehicle query method in this embodiment. For example... Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 and a memory 104 for storing data are also included. The processor 102 may be, but is not limited to, a microprocessor (MCU) or a programmable logic device (FPGA). The terminal may also include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that… Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown are illustrated.
[0046] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the vehicle query method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, thereby implementing the aforementioned method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0047] The transmission device 106 is used to receive or send data via a network. This network includes a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 can be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0048] This embodiment provides a vehicle query method. Figure 2 This is a flowchart of the vehicle query method in this embodiment, such as... Figure 2 As shown, the process includes the following steps:
[0049] Step S210: Obtain the image of the target vehicle to be queried.
[0050] The vehicle query method of this embodiment can run on a server with data computing and processing capabilities within an intelligent transportation system. The target vehicle image to be queried can be a user-provided image that needs to be searched and matched against video data collected by the multi-camera monitoring network of the intelligent transportation system. For example, a user may provide one or more target vehicle images, hoping to determine the destination or status of the vehicle in the image, or hoping to find vehicles with similar models, or hoping to obtain the vehicle's trajectory over a certain time period. The user can upload the target vehicle image to the server through a human-computer interaction interface, or directly transmit the target vehicle image to the server via wired or wireless communication.
[0051] Step S220: Based on the trained reinforcement learning model, determine the filter combination for the vehicle video data; the filter combination includes filtering conditions for different dimensions.
[0052] Specifically, the vehicle video data can be video data collected in traffic scenarios (such as urban roads and highways) by a monitoring network composed of multiple cameras deployed in different locations within an intelligent transportation system. In some embodiments, the raw video stream can be preliminarily denoised and image quality enhanced to form the vehicle video data. Then, the vehicle video data can be divided into segments or frames and stored to form a video database, with each video frame labeled with necessary metadata, such as the unique identifier (ID) of the associated camera, the timestamp of the video, and the associated geographical location.
[0053] Filter combinations can include filtering conditions targeting different dimensions. For example, filtering conditions can be set based on dimensions such as vehicle color, shooting time period, vehicle type in the video, and vehicle speed range. Combining multiple sets of filtering conditions can form a filter combination capable of multi-level filtering of vehicle video data. Several filtering conditions can be pre-defined. Then, using a trained reinforcement learning model, in different query scenarios (such as finding a specific vehicle, finding similar vehicles, finding the trajectory of a vehicle of interest, etc.), vehicle video data is extracted, and the video characteristics of the extracted vehicle video data are analyzed to determine the optimal filter combination for the vehicle video data. This achieves a balance between query efficiency and accuracy when querying vehicles.
[0054] The reinforcement learning model can be trained offline. When training the reinforcement learning model, several training video datasets with different video characteristics can be collected as a training set to train the model, enabling it to determine the optimal filter combination for different video characteristics.
[0055] Step S230: Based on the filter combination, the vehicle video data is filtered at least according to the target vehicle image to obtain filtered video data related to the target vehicle image.
[0056] After determining the optimal filter combination, vehicle video data can be filtered based on the filtering conditions set by that combination, according to the specific features of the target vehicle image corresponding to each filtering condition in the filter combination. For example, if the optimal filter combination is determined to be vehicle color, vehicle type, and shooting scene, and the target vehicle image shows a white vehicle, a highway shooting scene, and a small passenger car type, then data where the vehicle color is not white, the vehicle type is not a small passenger car, and the shooting scene is not a highway can be filtered out from the vehicle video data, retaining only the vehicle video data that matches the filtering conditions. Using filter combinations to filter vehicle video data collected by multiple cameras can improve the accuracy and efficiency of subsequent queries by reducing interference from irrelevant data, and on the other hand, it can also reduce the amount of data that needs to be processed in subsequent queries, thereby reducing the overall computational burden.
[0057] Step S240: Detect vehicles in the filtered video data, extract vehicle features from the detected vehicle targets, and construct a feature library based on the vehicle features.
[0058] Vehicle targets can be extracted from filtered video data using object detection, and their global or local features, such as vehicle color, license plate features, and shape features, can be extracted to form feature vectors. Then, all vehicle target feature vectors are stored uniformly in a feature library, and corresponding metadata such as vehicle ID, camera ID, and frame number can be maintained for easy subsequent retrieval. It is understood that this embodiment does not specifically limit the type of features extracted.
[0059] Step S250: Match the target vehicle features in the target vehicle image with the vehicle features in the feature library, and determine the candidate vehicle list that matches the target vehicle image based on the matching results.
[0060] This process involves pre-extracting corresponding vehicle features from the target vehicle image, such as vehicle color, license plate features, and shape features, as target vehicle features. Then, using a pre-defined similarity comparison method, such as Euclidean distance (L2 distance) or cosine similarity, the target vehicle features are compared with vehicle features in a feature library to determine the similarity between each vehicle target in the filtered video data and the vehicle targets in the target vehicle image. The similarity ranking is then output, and the top-ranked entries form a candidate vehicle list for matching the target vehicle image.
[0061] In related technologies, a monitoring network composed of multiple cameras is often used to collect video data in traffic scenarios, and then the video data collected by multiple cameras is directly processed. Specifically, for the video stream captured by each camera, traditional or deep learning algorithms are used for vehicle detection. Based on the detection results, vehicle trajectories are tracked in consecutive video frames, and tracking IDs are assigned. Features of vehicles observed from different cameras are compared to determine if they belong to the same vehicle entity, and based on the determination results, video data belonging to the same vehicle entity are fused. This approach leads to large-scale data processing overhead. Traditional multi-camera tracking often requires full detection and tracking of continuous video streams. As the number of cameras increases or the video resolution improves, the computational and storage costs of the data also increase dramatically. Furthermore, offline queries of massive video data in related technologies usually rely on manually setting single filter conditions, which cannot adapt to real-time or near real-time query requirements, and the retrieval accuracy is low, making it difficult to balance query efficiency and accuracy.
[0062] This embodiment introduces an intelligent combination of multiple filters and reinforcement learning models, which can filter irrelevant data before object detection processing of video data, thereby reducing the amount of data that needs high-precision processing and reducing the overall computational overhead; in addition, the reinforcement learning model can also balance query accuracy and efficiency in the selection of filter combinations.
[0063] Therefore, through steps S210 to S250, the target vehicle image to be queried is obtained; based on the trained reinforcement learning model, a filter combination for the vehicle video data is determined; the filter combination includes filtering conditions for different dimensions; based on the filter combination, the vehicle video data is filtered at least according to the target vehicle image to obtain filtered video data related to the target vehicle image; vehicle detection is performed on the filtered video data, vehicle features are extracted from the detected vehicle targets, and a feature library is constructed based on the vehicle features; the target vehicle features in the target vehicle image are matched with each vehicle feature in the feature library, and a candidate vehicle list matching the target vehicle image is determined based on the matching results. This process filters vehicle videos collected by multiple camera devices, thereby reducing the amount of data required for subsequent target detection and reducing computational overhead.
[0064] In one embodiment, based on step S220 above, determining the filter combination for the vehicle video data according to the trained reinforcement learning model may include:
[0065] Based on the trained reinforcement learning model, corresponding filter combinations are determined for each video segment in the vehicle video data according to the changes in video characteristics.
[0066] Vehicle video data collected by multiple cameras can first be divided into several video segments. Different filter combinations can be selected for different video segments to achieve frame-level filtering. Specifically, to achieve more efficient and accurate filtering for video segments with different video characteristics, a pre-trained reinforcement learning model can be used to extract vehicle video data. During the extraction process, appropriate filter combinations are selected based on the video characteristics of the vehicle video data. These video characteristics can characterize the image features of the video frames, such as scene, color, texture, and viewpoint. The appropriate filter combination can be selected based on the detectable video characteristics in the vehicle video data. For example, if a segment of vehicle video data has significant color contrast and rich texture, the reinforcement learning model is more likely to select a filter combination that combines color and texture features to improve filtering efficiency. Furthermore, after selecting a filter combination, real-time feedback from the filtering process, such as changes in video characteristics (e.g., scene changes), can be used to determine whether to change the filter combination to achieve faster filtering speed and higher accuracy.
[0067] In this embodiment, by utilizing the trained reinforcement learning model, it is possible to assign corresponding filter combinations to different video segments based on the video characteristics of vehicle video data to achieve frame-level data filtering. This enables more accurate and fine-grained data filtering in the pre-stage, reducing the amount of data that needs to be processed for subsequent target detection, thereby reducing the overall computational burden.
[0068] In another embodiment, based on step S230 above, and based on the filter combination, the vehicle video data is filtered at least according to the target vehicle image to obtain filtered video data related to the target vehicle image. Specifically, this may include:
[0069] Based on filter combinations, vehicle video data is filtered according to the target vehicle image and set query conditions to obtain filtered video data related to the target vehicle image.
[0070] During the process of reading vehicle video data, video frames irrelevant to the target vehicle image can be filtered out based on various filtering conditions included in the filter combination, while retaining video frames with high relevance, thus obtaining filtered video data related to the target vehicle image. In addition to providing the target vehicle image, users can also attach query conditions in the form of text descriptions, such as specifying time range constraints in textual descriptions. For example, if the filter combination is based on color features and vehicle type, and the target vehicle image shows a white vehicle and a small passenger car type, and the user's query condition is 9:00 AM to 10:00 AM, then during filtering, only video frames in the vehicle video data that are white, of the small passenger car type, and appear between 9:00 AM and 10:00 AM will be retained as filtered video data.
[0071] This embodiment can provide more flexible and diversified video data filtering based on the combination of target vehicle images and query conditions.
[0072] Furthermore, in one embodiment, vehicle detection is performed on the filtered video data, vehicle features are extracted from the detected vehicle targets, and a feature library is constructed based on the vehicle features. This may include:
[0073] Vehicles are detected in the filtered video data according to the preset target detection model to obtain vehicle targets; vehicle features are extracted from the vehicle targets according to the preset target extraction model.
[0074] Specifically, to achieve efficient and more versatile target detection, lightweight detection models can be employed. For example, the YOLOv8n model can be used to detect vehicle targets in filtered video data. Re-ID models can be used to extract global or local features of the vehicle targets, such as color, license plate features, and shape. This embodiment can perform sampling based on a lightweight target detection model, thereby achieving efficient vehicle detection with limited hardware resources.
[0075] In one embodiment, performing vehicle detection on the filtered video data according to a preset target detection model to obtain vehicle targets may include:
[0076] Based on the object detector interface, vehicle targets are detected in filtered video data using a pre-selected object detection model. The open object detector interface allows users to customize object detection model inputs, facilitating model replacement based on the scenario. Furthermore, the Re-ID feature library can be flexibly expanded to meet the needs of larger-scale urban surveillance or special scenarios. Thus, this embodiment achieves a reusable and scalable detection mechanism, adapting to the detection requirements of different scenarios.
[0077] In another embodiment, the vehicle query method described above may further include:
[0078] The system tracks vehicle targets detected in filtered video data to obtain tracking information. This can be achieved through offline ID tracking of the filtered video data based on a pre-defined target tracking model, outputting tracking information such as frame number, tracking ID, vehicle category, and bounding box. For example, YOLOv8n combined with the high-performance multi-target tracking model ByteTrack can be used to perform offline vehicle detection and tracking. A unique tracking ID is generated for each vehicle target, and the corresponding frame number and camera information are recorded to facilitate tracing the time and location information of vehicle appearance.
[0079] In one embodiment, after obtaining the candidate vehicle list, the vehicle query method may further include:
[0080] Based on the tracking information, the spatiotemporal information of each vehicle target in the candidate vehicle list is determined. After obtaining the candidate vehicle list that matches the target vehicle image, the time and location of the corresponding vehicle target appearing in the associated camera can be determined based on the tracking ID, frame number, and camera information contained in the tracking information, thereby determining the spatiotemporal information of each vehicle target.
[0081] In another embodiment, vehicle detection is performed on the filtered video data, vehicle features are extracted from the detected vehicle targets, and a feature library is constructed based on the vehicle features. This may include:
[0082] All vehicle features in the feature library are stored in a pre-defined vector retrieval index structure. Lightweight Re-ID models, such as MobileNet or more complex architectures, can be used to extract feature vectors of vehicle targets as vehicle features and store them in the feature library. Then, a vector retrieval index structure, such as the nearest neighbor search structure Faiss.IndexFlatL2, is used to build a vector index for the Re-ID feature library.
[0083] In this embodiment, considering that the feature library of all vehicle targets is a large storage unit, setting an index can make queries faster, improve the efficiency of target queries, and ensure query accuracy.
[0084] The most basic index building method, Faiss.IndexFlatL2, can be chosen to reduce the cost of index building while ensuring the usability of the constructed index. Understandably, other index building methods can also be used, but a method must be chosen that balances construction cost and query performance. This embodiment uses a vector retrieval index structure to build an index for the feature library, facilitating efficient similarity retrieval and improving the efficiency of multi-camera vehicle queries. Furthermore, by combining adjustable similarity thresholds and recall rates, an optimal balance can be found between efficiency and precision, thereby achieving a flexible indexing and retrieval mechanism.
[0085] In one embodiment, matching the target vehicle features in the target vehicle image with vehicle features in a feature library, and determining a candidate vehicle list matching the target vehicle image based on the matching results, may include:
[0086] The query method of the vector retrieval index structure is invoked to match the stored vehicle features against the target vehicle features, obtaining the matching results. The vector retrieval index structure encapsulates callable query methods. Therefore, when it is necessary to query vehicle features in the feature library and compare their similarity with the target vehicle features, the query method of the vector retrieval index structure can be directly invoked. For example, the query (search) method of Faiss.IndexFlatL2 can be directly invoked to calculate the Euclidean distance between the feature vector of the target vehicle feature and the feature vectors of all vehicle features in the feature library, and return the K results with the smallest Euclidean distance. After sorting by similarity, the most matching vehicle entries can be output to the user. Therefore, this embodiment can achieve efficient vehicle feature querying and comparison based on the vector retrieval index structure, and can obtain a list of suspected vehicles based on similarity thresholds and recall control, displaying the time range in which each suspected vehicle appears on different cameras.
[0087] Figure 3 This is a schematic diagram illustrating the training and application of a reinforcement learning agent according to this embodiment. The reinforcement learning agent may include filter combination and a reinforcement learning model. Figure 3 As shown, during offline training of the reinforcement learning model, training video data is collected as training samples. The training video data is divided into training query images. Several feature extractors are used to extract features from each frame of the query image, and the filtered video data is obtained. Target tracking is performed on the training video data to obtain rectangular bounding boxes, and tracking IDs are recorded, such as ID1, ID2, ID3, ..., IDn. The reinforcement learning agent is trained based on the tracking results and the filtered video data (ID1 video, ID2 video, IDn video). During online queries, the query video data is input into the reinforcement learning agent for filtering based on the user-input query image. This query video data is the vehicle video data collected by multiple cameras, and the query image is the target vehicle image. Target tracking is performed on the filtered video data, outputting rectangular bounding boxes. A Re-ID model is used for feature extraction to construct a feature library. After feature extraction from the query image using the Re-ID model, the feature vector of the query image is compared with the feature vectors in the feature library, and then sorted according to similarity to obtain the final candidate set. The candidate set represents a list of candidate vehicles that match the query image.
[0088] Figure 4 Here are flowcharts of some embodiments of the vehicle query method, such as Figure 4 As shown, the vehicle query method includes the following steps:
[0089] Step S401: Receive the target vehicle image uploaded by the user.
[0090] Step S402: Extract features from the target vehicle image based on the target detector and Re-ID model to obtain the target vehicle features.
[0091] Step S403: Acquire vehicle video data; the vehicle video data is collected by multiple cameras deployed at different locations on urban roads.
[0092] Step S404: Use the trained reinforcement learning model to determine the filter combination for different video segments of the vehicle video data.
[0093] Step S405: Using filter combination, the vehicle video data is processed according to the target vehicle image and the query conditions input by the user to obtain filtered video data.
[0094] Step S406: The filtered video data is processed by using YOLOv8n in conjunction with ByteTrack to perform offline vehicle detection and tracking. A unique tracking ID is generated for each detected vehicle target, and the corresponding frame number and camera information are recorded.
[0095] Step S407: Use Faiss.IndexFlatL2 to store the feature vector of each vehicle target.
[0096] Step S408: Call the query method of Faiss.IndexFlatL2 to match and sort the similarity of the target vehicle features with the stored vehicle features.
[0097] Step S409: Based on the similarity threshold and recall control, obtain a list of candidate vehicles that match the target vehicle image, and display the time range in which each candidate vehicle appears in different cameras.
[0098] Steps S401 to S409 above, through the combination of multiple filters and reinforcement learning models, can reduce the amount of data requiring high-precision processing and lower the overall computational cost. By employing flexible indexing and retrieval mechanisms, the efficiency of multi-camera vehicle queries can be improved, balancing efficiency and accuracy. Therefore, it can enhance efficiency and accuracy in multi-camera vehicle tracking and query scenarios, and provide more comprehensive and intelligent support for urban traffic management and safety monitoring.
[0099] This embodiment also provides a vehicle query device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. The terms "module," "unit," "subunit," etc., used below refer to combinations of software and / or hardware that implement a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0100] Figure 5This is a structural block diagram of the vehicle query device 50 in this embodiment, as shown below. Figure 5 As shown, the vehicle query device 50 includes: an acquisition module 51, a determination module 52, a filtering module 53, a detection and extraction module 54, and a matching module 55; wherein:
[0101] The module 51 is used to acquire the target vehicle image to be queried; the module 52 is used to determine the filter combination for the vehicle video data based on the trained reinforcement learning model; the filter module 53 is used to filter the vehicle video data based on the filter combination, at least according to the target vehicle image, to obtain filtered video data related to the target vehicle image; the detection and extraction module 54 is used to perform vehicle detection on the filtered video data, extract vehicle features from the detected vehicle targets, and construct a feature library based on the vehicle features; the matching module 55 is used to match the target vehicle features in the target vehicle image with each vehicle feature in the feature library, and determine a list of candidate vehicles that match the target vehicle image based on the matching results.
[0102] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.
[0103] In one embodiment, the determining module 52 is used to determine corresponding filter combinations for each video segment in the vehicle video data based on the changes in video characteristics of the vehicle video data after training the reinforcement learning model.
[0104] In one embodiment, the filtering module 53 is used to filter vehicle video data based on a filter combination, according to the target vehicle image and set query conditions, to obtain filtered video data related to the target vehicle image.
[0105] In one embodiment, the detection and extraction module 54 includes a target detection unit and a target extraction unit; wherein, the target detection unit is used to perform vehicle detection on the filtered video data according to a preset target detection model to obtain vehicle targets; and the target extraction unit is used to extract vehicle features from the vehicle targets according to a preset target extraction model.
[0106] In one embodiment, the target detection unit is used to detect vehicles in filtered video data using a pre-selected target detection model based on the target detector interface, thereby obtaining vehicle targets.
[0107] In one embodiment, the vehicle query device further includes a tracking module, which is used to track vehicle targets detected in filtered video data to obtain tracking information.
[0108] In one embodiment, the matching module 55 is further configured to determine the spatiotemporal information of each vehicle target in the candidate vehicle list based on the tracking information.
[0109] In one embodiment, the detection and extraction module 54 is used to store all vehicle features in the feature library into a preset vector retrieval index structure.
[0110] In one embodiment, the matching module 55 is used to call the query method of the vector retrieval index structure to match the stored vehicle features according to the target vehicle features and obtain the matching result.
[0111] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated in this embodiment.
[0112] It should be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. All other embodiments derived by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.
[0113] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0114] Obviously, the accompanying drawings are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar situations based on these drawings without any creative effort. Furthermore, it is understood that although the work done in this development process may be complex and lengthy, for those skilled in the art, certain design, manufacturing, or production modifications made based on the technical content disclosed in this application are merely conventional technical means and should not be considered as insufficient disclosure of this application.
[0115] The term "embodiment" in this application refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply that it is mutually exclusive with or independent of other embodiments. It will be clearly or implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0116] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.
Claims
1. A vehicle query method, characterized in that, include: Obtain the image of the target vehicle to be queried; Based on the trained reinforcement learning model, determine the filter combination for vehicle video data; The filter combination includes filtering conditions for different dimensions; Based on the filter combination, the vehicle video data is filtered at least according to the target vehicle image to obtain filtered video data related to the target vehicle image; Vehicle detection is performed on the filtered video data, vehicle features are extracted from the detected vehicle targets, and a feature library is constructed based on the vehicle features; The target vehicle features in the target vehicle image are matched with the vehicle features in the feature library, and a list of candidate vehicles that match the target vehicle image is determined based on the matching results.
2. The vehicle query method according to claim 1, characterized in that, Based on the trained reinforcement learning model, a combination of filters for the vehicle video data is determined, including: Based on the trained reinforcement learning model, and according to the changes in the video characteristics of the vehicle video data, a corresponding filter combination is determined for each video segment in the vehicle video data.
3. The vehicle query method according to claim 1, characterized in that, Based on the filter combination, the vehicle video data is filtered at least according to the target vehicle image to obtain filtered video data related to the target vehicle image, including: Based on the filter combination, the vehicle video data is filtered according to the target vehicle image and the set query conditions to obtain filtered video data related to the target vehicle image.
4. The vehicle query method according to claim 1, characterized in that, Vehicle detection is performed on the filtered video data, vehicle features are extracted from the detected vehicle targets, and a feature library is constructed based on the vehicle features, including: Vehicle targets are obtained by detecting vehicles in the filtered video data according to a preset target detection model. Based on a preset target extraction model, vehicle features are extracted from the vehicle target.
5. The vehicle query method according to claim 4, characterized in that, Vehicle targets are obtained by performing vehicle detection on the filtered video data according to a preset target detection model, including: Based on the target detector interface, vehicle targets are obtained by using a pre-selected target detection model to detect vehicles in the filtered video data.
6. The vehicle query method according to claim 1, characterized in that, The method further includes: The vehicle targets detected in the filtered video data are tracked to obtain tracking information.
7. The vehicle query method according to claim 6, characterized in that, After obtaining the candidate vehicle list, the method further includes: Based on the tracking information, the spatiotemporal information of each vehicle target in the candidate vehicle list is determined.
8. The vehicle query method according to claim 1, characterized in that, Vehicle detection is performed on the filtered video data, vehicle features are extracted from the detected vehicle targets, and a feature library is constructed based on the vehicle features, including: All vehicle features in the feature library are stored in a preset vector retrieval index structure.
9. The vehicle query method according to claim 8, characterized in that, The target vehicle features in the target vehicle image are matched with vehicle features in the feature library. Based on the matching results, a list of candidate vehicles matching the target vehicle image is determined, including: The query method of the vector retrieval index structure is invoked to match the stored vehicle features according to the target vehicle features, and the matching results are obtained.
10. A vehicle query device, characterized in that, include: The module comprises an acquisition module, a determination module, a filtering module, a detection and extraction module, and a matching module; among which: The acquisition module is used to acquire the image of the target vehicle to be queried; The determining module is used to determine the filter combination for vehicle video data based on the trained reinforcement learning model. The filtering module is used to filter the vehicle video data based on the filter combination, at least according to the target vehicle image, to obtain filtered video data related to the target vehicle image; The detection and extraction module is used to detect vehicles in the filtered video data, extract vehicle features from the detected vehicle targets, and construct a feature library based on the vehicle features. The matching module is used to match the target vehicle features in the target vehicle image with the vehicle features in the feature library, and determine a list of candidate vehicles that match the target vehicle image based on the matching results.