Vehicle-mounted camera data automatic processing method

By automatically processing vehicle camera data through a pre-trained model, extracting image and text features, and combining metadata for comprehensive similarity calculation, the problem of low efficiency and manual intervention in traditional image databases is solved, achieving efficient and accurate image retrieval and management.

CN120997788APending Publication Date: 2025-11-21ANHUI JIANGHUAI AUTOMOBILE GRP CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511177890.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Traditional image databases are inefficient and lack retrieval accuracy in the field of intelligent assisted driving. They cannot understand the semantic relationships described by natural language, making it difficult to meet the needs of efficient processing of large-scale data. Furthermore, the need for manual intervention in the transmission of image data from vehicles leads to high management costs and response delays.

Method used

It employs pre-trained visual detection and language understanding models to automatically process vehicle camera data, extract image visual and semantic features, and perform comprehensive similarity calculations in conjunction with metadata, thereby enabling image retrieval under natural language descriptions and supporting zero-shot learning in new scenarios.

Benefits of technology

It achieves fully automated data management without human intervention, improves the efficiency and accuracy of data collection and storage, lowers the query threshold, supports users to accurately locate images through natural language descriptions, adapts to new scenario retrieval, and reduces management costs and response latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997788A_ABST
    Figure CN120997788A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle-mounted camera data automatic processing method, which has the main conception that image data returned by a vehicle is automatically collected, and a query text input by a user is received; extracting image visual features of the preprocessed image data by using a pre-trained visual detection model; utilizing a pre-trained language understanding model to extract semantic features of the query text; combining the image visual features and the semantic features to obtain comprehensive similarity; and obtaining a target image with the highest matching degree with the query text from the image data according to a sorting result of the comprehensive similarity. Semantic association of the image and the text is realized through pre-training, zero sample learning of a new scene is supported, an unlabeled new scene image can be accurately retrieved without fine tuning, and the problems that a traditional scheme cannot accurately position the image according to user description, depends on specific data fine tuning, cannot adapt to the new scene, and cannot accurately retrieve the unlabeled new scene image are effectively solved. And response is delayed and management cost is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of driver assistance technology, and in particular to an automatic data processing method for vehicle-mounted cameras. Background Technology

[0002] With the rapid development of fields such as intelligent assisted driving and intelligent transportation, the amount of image data transmitted back by vehicles is growing exponentially. Traditional image databases rely on manual annotation or structured queries, which suffer from low efficiency and insufficient retrieval accuracy, making it difficult to meet the needs of efficient processing of large-scale data.

[0003] In this process, images are manually labeled with simple tags such as "daytime," "rainy day," "cat," and "dog." Image feature extraction algorithms are then used to generate feature vectors for these images. These feature vectors are used to calculate image similarity, thus enabling content-based image retrieval. Traditional image database retrieval relies primarily on keyword matching and cannot understand the semantic relationships described in natural language. Users need to accurately input keywords to find relevant images, which is insufficient for fuzzy or complex queries (such as "multiple vehicles converging in a tunnel").

[0004] In addition, some industry experts have proposed a solution using BERT + object detection, where BERT is used to recognize text and object detection is responsible for recognizing images. For example, given an input image (a red car driving in the rain), the output description is "a red car driving in the rain." BERT parses the text (such as the semantics of "driving in the rain"), and YOLO detects the "red car" in the image. However, after the features from BERT and object detection are extracted independently, they are fused through simple concatenation or attention mechanisms, making it difficult to capture deep cross-modal relationships. Furthermore, this solution requires large-scale annotation of both image and text data separately. Summary of the Invention

[0005] In view of the above, the present invention aims to provide an automatic data processing method for vehicle-mounted cameras to solve the aforementioned technical problems.

[0006] The technical solution adopted in this invention is as follows:

[0007] This invention provides an automatic data processing method for vehicle-mounted cameras, comprising:

[0008] Automatically collect image data transmitted back by vehicles and receive query text input by users;

[0009] The image visual features of the preprocessed image data are extracted using a pre-trained visual detection model;

[0010] The semantic features of the query text are extracted using a pre-trained language understanding model;

[0011] By combining the image visual features and the semantic features, a comprehensive similarity is obtained;

[0012] Based on the sorting results of the comprehensive similarity, the target image with the highest matching degree with the query text is obtained from the image data.

[0013] In at least one of the possible implementations, after returning the target image, if it is determined that the user does not accept the current target image, the query text is automatically adjusted and alternative query information is generated to allow the user to supplement the query conditions.

[0014] In at least one of the possible implementations, the automatic acquisition of image data transmitted back by the vehicle includes: automatically receiving and storing the original image data transmitted back by the vehicle to the cloud or a local database, while recording at least the following metadata: timestamp and location information.

[0015] In at least one possible implementation, extracting the image visual features specifically includes:

[0016] The key targets in the image data were identified using a pre-trained visual detection model;

[0017] Extract the corresponding image visual features from the key targets;

[0018] The image visual features are combined with the metadata and stored in a structured manner to automatically obtain image tags.

[0019] In at least one possible implementation, extracting the semantic features of the query text includes: extracting the overall semantic features of the query text using a pre-trained language understanding model, and generalizing the query text and encoding it into a vector.

[0020] In at least one of the possible implementations, a visual detection model and a language understanding model are pre-trained using pre-extracted image-text pairs related to vehicles and traffic.

[0021] Compared with existing technologies, the main design concept of this invention lies in providing a fully automated solution from data acquisition, storage, preprocessing to retrieval feedback. No manual intervention is required; users can describe their needs using natural language, and the system can automatically parse semantics and match image data. Specifically, it automatically receives and stores raw image data transmitted from vehicles, along with associated sensor data, time, location, and other metadata, achieving full lifecycle management of data and significantly improving the efficiency and accuracy of data acquisition and storage. In practical applications, users can describe their query needs using natural language, and the system automatically extracts overall semantic features, achieving accurate understanding of user intent, lowering the query threshold, and enhancing the user experience. Furthermore, it returns the image or video clip with the highest matching degree, supporting users to download it from the cloud or locate the video image themselves based on the returned video recording time point, facilitating the visualization and interactive viewing of search results.

[0022] This invention effectively solves the problems of traditional solutions failing to meet users' needs for accurate image location based on text descriptions, relying on specific data for fine-tuning, and being unable to adapt to new scenarios. This invention achieves semantic association between images and text through pre-training, supports zero-shot learning of new scenarios, and can accurately retrieve unlabeled images of new scenarios (such as unlabeled data retrieval for sudden accident scenarios) without fine-tuning. Furthermore, this invention also solves the problem of requiring manual intervention and preprocessing (such as noise reduction and annotation) of vehicle-transmitted image data, leading to response delays and high management costs. Attached Figure Description

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described below with reference to the accompanying drawings, wherein:

[0024] Figure 1 This is a schematic diagram of an automatic data processing method for vehicle-mounted cameras provided in an embodiment of the present invention. Detailed Implementation

[0025] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0026] This invention proposes an embodiment of an automatic data processing method for vehicle-mounted cameras, specifically, as follows: Figure 1 As shown, it includes:

[0027] Step S1: Automatically collect image data transmitted back by the vehicle and receive query text input by the user;

[0028] Specifically, the system automatically receives and stores the raw image data transmitted back by the vehicle to the cloud or local database, while recording metadata such as sensor data, timestamps, and location (GPS coordinates).

[0029] Step S2: Extract the image visual features of the pre-processed image data using a pre-trained visual detection model;

[0030] The preprocessing mentioned here may include, but is not limited to, image denoising, contrast enhancement, and resolution adjustment, in order to improve the accuracy of subsequent analysis.

[0031] Regarding the pre-training process, image-text pairs related to traffic and vehicle images can be extracted from the Internet to pre-train a visual detection model and a language understanding model (in practice, the two models can be understood as a joint model with different tasks). Specific training data sources can include web page images and their descriptions, social media content, etc., so that the model can learn rich visual and language concepts.

[0032] Continuing from the previous section, a pre-trained visual detection model is used to identify key targets in the image data, such as vehicles, roads, and weather features (e.g., raindrops, fog). Corresponding visual features (e.g., vehicle position, distance, weather markers) are extracted from these features and further structured and stored in conjunction with the metadata mentioned earlier (e.g., "distance between the vehicle in front and the vehicle behind = 8.5 meters", "raindrop detection confidence = 92%)", which serve as image labels.

[0033] Step S3: Extract the semantic features of the query text using a pre-trained language understanding model;

[0034] To elaborate, after obtaining the text description submitted by the user (such as "a scenario of vehicles tailgating on a highway in rainy weather"), the pre-trained language understanding model mentioned above is used to extract the overall semantic features. At the same time, synonyms and hyponyms are generated (such as "rainy day" → "rainfall" and "heavy rain"), and the query text is encoded into a vector.

[0035] Step S4: Combine the visual features of the image with the semantic features to obtain the comprehensive similarity;

[0036] The semantic features and image visual features are compared to calculate their similarity. Then, the timestamps and GPS coordinates in the metadata mentioned above are used to filter out images that do not meet the criteria.

[0037] Step S5: According to the sorting result of the comprehensive similarity, obtain the target image with the highest matching degree with the query text from the image data.

[0038] Images are sorted by comprehensive similarity (text + visual), and the image data with the highest matching degree is automatically returned (the image is not limited to static or dynamic, for example, a picture of "two cars 7 meters apart in a rainy scene" is retrieved from the image data based on similarity and returned as the text result of "rainy highway vehicle tailgating scene"). Of course, better yet, in addition to automatic sorting by similarity, users can also pre-define other sorting rules (such as time priority, distance priority, etc.) and can also make other settings for the returned images, such as requiring image thumbnails, key information annotations (such as "raindrop confidence 92%"), timeline positioning, etc.

[0039] Furthermore, the returned final image results can support cloud download of static images and video clips (with timestamps), or can be triggered to jump directly to a specified time point in the original video stream.

[0040] Furthermore, based on the returned target image, a question can be asked to the user whether they accept the result. If the user is not satisfied with the current result, the query text can be automatically adjusted and alternative query information can be generated (such as "Do you need to add the 'nighttime' condition?"). Alternatively, the user can be asked to readjust the query input, or the top N results with comprehensive similarity can be provided, from which the user can mark the results they accept. Based on the user's selection, the aforementioned algorithm model can also be actively learned to update the matching processing strategy.

[0041] In summary, the main concept of this invention is to automatically collect image data transmitted from vehicles and receive query text input by users; extract image visual features from the pre-processed image data using a pre-trained visual detection model; extract semantic features from the query text using a pre-trained language understanding model; combine the image visual features and the semantic features to obtain a comprehensive similarity score; and obtain the target image with the highest matching degree to the query text from the image data according to the ranking result of the comprehensive similarity score. This invention achieves semantic association between images and text through pre-training, supports zero-shot learning of new scenes, and can accurately retrieve unlabeled new scene images without fine-tuning. It effectively solves a series of problems associated with traditional solutions, such as the inability to accurately locate images based on user descriptions, reliance on specific data for fine-tuning, inability to adapt to new scenes, response delays, and high management costs.

[0042] In this invention, when directional terms are mentioned, they are relative concepts based on the embodiments. Furthermore, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, A and B simultaneously, or B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects have an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0043] The above description of the structure, features, and effects of the present invention is based on the embodiments shown in the figures. However, the above are only preferred embodiments of the present invention. It should be noted that the technical features involved in the above embodiments and their preferred methods can be reasonably combined and matched by those skilled in the art to form a variety of equivalent solutions without departing from or changing the design concept and technical effects of the present invention. Therefore, the present invention is not limited to the scope of implementation shown in the figures. Any changes made in accordance with the concept of the present invention, or modifications to equivalent embodiments, that do not exceed the spirit covered by the specification and figures, should be within the protection scope of the present invention.

Claims

1. A method for automatically processing data from a vehicle-mounted camera, characterized in that, include: Automatically collect image data transmitted back by vehicles and receive query text input by users; The image visual features of the preprocessed image data are extracted using a pre-trained visual detection model; The semantic features of the query text are extracted using a pre-trained language understanding model; By combining the image visual features and the semantic features, a comprehensive similarity is obtained; Based on the sorting results of the comprehensive similarity, the target image with the highest matching degree with the query text is obtained from the image data.

2. The automatic data processing method for vehicle-mounted cameras according to claim 1, characterized in that, After returning the target image, if it is determined that the user does not accept the current target image, the query text is automatically adjusted and alternative query information is generated to allow the user to supplement the query conditions.

3. The automatic data processing method for vehicle-mounted cameras according to claim 1, characterized in that, The automatic acquisition of image data transmitted from vehicles includes: automatically receiving and storing the original image data transmitted from vehicles to the cloud or local database, while recording at least the following metadata: timestamp and location information.

4. The automatic data processing method for vehicle-mounted cameras according to claim 3, characterized in that, Extracting the visual features of the image specifically includes: The key targets in the image data were identified using a pre-trained visual detection model; Extract the corresponding image visual features from the key targets; The image visual features are combined with the metadata and stored in a structured manner to automatically obtain image tags.

5. The automatic data processing method for vehicle-mounted cameras according to claim 1, characterized in that, The extraction of semantic features from the query text includes: using a pre-trained language understanding model to extract the overall semantic features from the query text, and then generalizing the query text and encoding it into a vector.

6. The automatic data processing method for vehicle-mounted cameras according to any one of claims 1 to 5, characterized in that, The visual detection model and the language understanding model are pre-trained using pre-extracted image-text pairs related to vehicles and traffic.