Image query method, vehicle and electronic equipment
Through feature encoding and database query strategies for images in open scenes, the problem of low image query accuracy in unmanned driving is solved, and the effect of improving image query accuracy and efficiency is achieved.
Patent Information
- Application Number
- CN202510664948.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
AI Technical Summary
In the open scene of unmanned driving, the image screening and recognition methods lack effective discriminant models, resulting in the need of manpower or inefficient algorithms to check each image one by one, which consumes a lot of resources and lacks accuracy, resulting in low image query accuracy.
By obtaining the input data to be queried, feature encoding is performed, the initial feature is obtained, and the target features matching the initial feature is queryed in the original image sample features and the foreground image sample features of the database, and the target image is output or labeled.
It realizes efficient and accurate identification of images matching the target object to be query in open scenes, significantly improving the accuracy and efficiency of image query, reducing interference with irrelevant background information, and improving sensitivity to scarce targets.
Smart Images

Figure CN120179849A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of driverless technology and computer technology. Specifically, the present invention relates to an image query method, a vehicle, and an electronic device. Background Art
[0002] Currently, in the context of industrial automation and intelligence, especially for driverless vehicles in open scenarios, image recognition and understanding technologies play a crucial role in achieving effective environmental perception, resource management, and automated operations. However, in open scenarios, the images captured by driverless vehicles often contain large areas of background and a small number of foreground targets, which poses significant challenges to image screening and recognition.
[0003] In related technologies, image screening and recognition methods usually rely on a basic model for a specific task as a discriminator to determine whether an image contains the required information. However, in cold start scenarios, the above methods often result in a huge workload because they cannot accurately identify and filter key information in images.
[0004] That is, in the open scenario of driverless driving, when using a driverless vehicle to recognize images captured in an open scenario, the above methods lack an effective discriminant model, which means that each image needs to be checked one by one through manual labor or inefficient algorithms to determine whether it contains foreground targets related to the task. This not only consumes a large amount of resources but also results in insufficient accuracy in image screening and recognition in open scenarios. Therefore, there is still a technical problem of low query accuracy for images.
[0005] For the above problems, no effective solution has been proposed yet. Summary of the Invention
[0006] Embodiments of the present invention provide an image query method, a vehicle, and an electronic device to at least solve the technical problem of low query accuracy for images.
[0007] According to one aspect of the embodiments of the present invention, an image query method is provided, including: obtaining input data to be queried, where the input data is used to represent a target object in an open scenario to be queried; encoding the input data to obtain initial features; querying, among the original image sample features corresponding to the original image samples in the database and the foreground image sample features corresponding to the foreground image samples, for target features that match the initial features, where there is a mapping relationship between the original image samples and the foreground image samples, the original image samples are used to represent image samples with an image of an open scenario sample as the background image sample and an image of an object sample as the foreground image sample, and the proportion of the size of the object sample in the original image sample is less than a set threshold; labeling or outputting the target images corresponding to the target features.
[0008] According to another aspect of the embodiments of the present invention, there is also provided an image query device, including: an acquisition unit configured to acquire input data to be queried, where the input data is used to represent a target object in an empty scene to be queried; an encoding unit configured to encode the input data to obtain an initial feature; a query unit configured to query, from the original image sample features corresponding to the original image samples in the database and the foreground image sample features corresponding to the foreground image samples, a target feature that matches the initial feature, where there is a mapping relationship between the original image samples and the foreground image samples, and the original image samples are used to represent image samples with an empty scene sample image as the background image sample and an object sample image as the foreground image sample, and the proportion of the size of the object sample in the original image sample is less than a set threshold; and a processing unit configured to label or output the target image corresponding to the target feature.
[0009] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium storing multiple instructions, and the instructions are suitable for being loaded and executed by a processor to execute any one of the above methods.
[0010] According to another aspect of the embodiments of the present invention, there is also provided an electronic device including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute any one of the above methods.
[0011] According to another aspect of the embodiments of the present invention, there is also provided a vehicle including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute any one of the above methods.
[0012] According to another aspect of the embodiments of the present invention, there is also provided a computer program product including a computer program, and the computer program can be used to execute any one of the above methods when executed by a processor.
[0013] In an embodiment of the present invention, input data of a target object in a to-be-query empty scene can be obtained, and the input data can be encoded to obtain an initial feature. A target feature matching the above initial feature can be queried from the original image sample features and foreground image sample features in a database. And a target image including the target feature and the target object can be output or labeled. In this embodiment, in the face of the challenge that the target object accounts for a small proportion in the empty scene, foreground image samples having a mapping relationship with the original image samples in the empty scene can be pre-identified and extracted, and a database can be established by combining the foreground image sample features and the original image sample features. In the query stage, the initial feature obtained by encoding the input data is efficiently compared with the original image sample features and the foreground image sample features in the database. Even in the cold start scenario lacking specific samples, an image matching the to-be-query target object can be accurately identified, significantly improving the accuracy and efficiency of image query in the empty scene. Through the above feature encoding of the input data and database query strategy, the interference of irrelevant background information is greatly reduced, and the sensitivity to rare targets is improved, thus ensuring the image query effect in the empty scene. The technical problem of low query accuracy of images is solved, and the technical effect of improving the query accuracy of images is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention. In the drawings:
[0015] Figure 1 is a flowchart of a method for querying an image according to an embodiment of the present invention;
[0016] Figure 2 is a flowchart of a method for searching for an image by text and searching for an image by image according to an embodiment of the present invention;
[0017] FIG. 3(a) is a schematic diagram of a problem image according to an embodiment of the present invention;
[0018] FIG. 3(b) is a schematic diagram of another problem image according to an embodiment of the present invention;
[0019] FIG. 3(c) is a schematic diagram of another problem image according to an embodiment of the present invention;
[0020] FIG. 3(d) is a schematic diagram of another problem image according to an embodiment of the present invention;
[0021] Figure 4 is a schematic diagram of a network structure for only one view according to an embodiment of the present invention;
[0022] FIG. 5(a) is a schematic diagram of a search result obtained by searching by text for an image according to an embodiment of the present invention;
[0023] FIG. 5(b) is a schematic diagram of another search result obtained by searching by text for an image according to an embodiment of the present invention;
[0024] FIG. 5(c) is a schematic diagram of another search result obtained by searching by text for an image according to an embodiment of the present invention;
[0025] FIG. 5(d) is a schematic diagram of another search result obtained by searching by text for an image according to an embodiment of the present invention;
[0026] FIG. 5(e) is a schematic diagram of another search result obtained by searching by text for an image according to an embodiment of the present invention;
[0027] FIG. 5(f) is a schematic diagram of another search result obtained by searching by text for an image according to an embodiment of the present invention;
[0028] FIG. 6(a) is a schematic diagram of an input image of an image to be searched by searching by image according to an embodiment of the present invention;
[0029] FIG. 6(b) is a schematic diagram of a search result obtained by searching by image based on the input image according to an embodiment of the present invention;
[0030] FIG. 6(c) is a schematic diagram of another search result obtained by searching by image based on the input image according to an embodiment of the present invention;
[0031] Figure 7 is a schematic structural diagram of an image query device according to an embodiment of the present invention;
[0032] Figure 8 is a schematic diagram of an electronic device for an image query method according to an embodiment of the present invention. Detailed implementation manners
[0033] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0034] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0035] Embodiment 1
[0036] According to an embodiment of the present invention, an embodiment of an image query method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0037] Figure 1 is a flowchart of an image query method according to an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:
[0038] Step S102, obtain the input data to be queried.
[0039] In the technical solution provided in step S102 of the above embodiment of the present invention, the input data can be used to represent the target object in the empty scene to be queried. The input data can be text or an image. That is, the input data can be the text or image input for describing the target object in the image to be queried, that is, the input text or the input image. The flexibility of the form of the input data means that the user can input a text description of the target object or directly provide an image example of the target object as the basis for the query.
[0040] It should be noted that the above form of the input data is only for illustrative purposes and is not specifically limited here. As long as it is a form of input data that can be used to describe the target object in the empty scene to be queried, it is within the protection scope of the embodiments of the present invention, and no further examples will be given here.
[0041] Optionally, an open scene may refer to an environment where the proportion of the target object in the background of the image is relatively small. For example, a mine scene, etc. That is to say, an open scene may refer to a scene where the proportion of the target object included in the background is small. The target object may be a specific object or individual that the user needs to locate or identify in the image of the open scene. For example, it may be a worker, water, or a vehicle in a mine scene, etc. In an open scene, the target object may appear very small relative to the entire image, resulting in a sparse distribution of the features of the target object in the image, increasing the difficulty of recognition and location.
[0042] It should be noted that the above-mentioned open scene and the target object in the open scene are only for illustrative purposes and are not specifically limited here. As long as the scene where the proportion of the target object to be queried is small in a certain scene can be the above-mentioned open scene, and the physical objects or people in the open scene can all be the target object, and no further examples will be given here. The embodiments of the present invention only take the mine scene as an optional example.
[0043] In this embodiment, if it is necessary to query an image containing a target object in an open scene, input data for describing the target object in the open scene to be queried can be obtained.
[0044] Optionally, when it is necessary to query a specific target object in an open scene, the above-mentioned target object can be described in text form.
[0045] For example, the user may want to find images in a mine scene that contain keywords such as "worker", "yellow truck", or "ponding water", etc. Then, obtain the text description of the input text "yellow truck in the mine scene", and relevant mine scene images containing yellow trucks can be searched in the database based on the above text description. The user inputs the text "human with white truck", and relevant mine scene images containing workers and white trucks can be found in the database based on the above text description. Take "ponding water" as the input text to find images related to ponding water in the mine scene for subsequent analysis.
[0046] It should be noted that the input data obtained by the above "searching images by text" method is manifested as a text description of the target object, which can be a specific object name, feature description, or a query condition defined by the user. In the embodiments of the present invention, the content required for the input data in text form is not specifically limited and can be preset according to the open scene to be queried, the target object, and the user's needs.
[0047] Optionally, in addition to the text description, the embodiments of the present invention also support the user to directly upload or specify an example image containing the target object as input data. The above "image search by image" method is applicable when the user cannot accurately describe the target object in words, or the target object has complex appearance features, etc.
[0048] For example, if the user uploads a picture of a mine scene, in which there is a staff member operating equipment, other mine scene images of staff members similar to the features in the above picture can be queried from the database based on the features in the above picture. If a picture of an abnormal situation in the mine is received, in which the side of a yellow truck is visible and there is accumulated water. The above image can be used as input data to search for images in the database that have a similar perspective of the truck side and accumulated water situation as the input data, providing data support for subsequent event analysis or safety inspections.
[0049] It should be noted that when using an image as input data, the data obtained is a two-dimensional (2D) image captured by a camera or other image acquisition device, which intuitively reflects the appearance of the target object and the surrounding environment. The features of the above picture, including the shape, color, pose of the target object, and the interaction with the environment, will be encoded and used to compare with the image features in the feature database, so as to locate the images in the database that are similar to the initial features.
[0050] In the embodiments of the present invention, whether the input data is obtained through text or image, the core is to provide an accurate description of the target object for subsequent feature encoding and database query. Through the above steps, it is possible to effectively focus on the target object, and even in an empty scene, accurate and efficient image query can be carried out, significantly improving the accuracy and efficiency of the query. The above method is not only applicable to queries initiated by users, but also applicable to the system to automatically detect and analyze specific changes or abnormal situations in the scene, thus providing strong technical support for applications such as mine automation and safety monitoring.
[0051] Step S104, encode the input data to obtain initial features.
[0052] In the technical solution provided in step S104 of the embodiment of the present invention above, the initial feature can be obtained by encoding the input data through a deep learning model, for example, a Contrastive Language Image Pre-training (CLIP) model. The initial feature can be the encoded feature of the target object in the input data, and can also be referred to as the encoded feature or encoded authentication. The initial feature can be a vector encoding in the feature space, and the vector encoding reflects the position of the initial feature in the feature space and the key attributes of the input data, including the visual features and semantic descriptions of the target object. By encoding the input data, the features of the target object in the input data can be transformed into the feature space, so that the similarity between vectors (the features of the image or the features of the text and the features of the image) can be measured by the distance between the vectors in the feature space. That is, the smaller the distance, the higher the similarity between the features of the image or the features of the text and the features of the image. Therefore, by encoding the input data into the initial feature, the embodiment of the present invention can efficiently find an image close to the initial feature in the database based on the above initial feature.
[0053] It should be noted that the above method and deep learning model for encoding the input data are only for illustrative purposes and are not specifically limited here. As long as it is a method and process for analyzing the initial feature describing the target object from the input data, it is within the protection scope of the embodiment of the present invention.
[0054] In this embodiment, after obtaining the input data to be queried, the input data can be encoded to obtain the initial feature.
[0055] Optionally, the CLIP model is used to perform deep learning encoding on the input data (input text or input image), so as to generate an initial feature that can reflect the characteristics of the target object. The core of the CLIP model lies in its ability to learn the similarity between text and images across modalities. Through the method of contrastive learning and training on a large-scale unlabeled dataset, it has learned how to map text descriptions and corresponding image contents to the same embedding space. In this embedding space, similar contents (whether text or image) are encoded as close vector representations, so that the CLIP model can query the database for matching images according to the input text or image.
[0056] For example, if the input data is text, the CLIP model can transform the input data into a series of word embeddings and process them through multiple layers of Transformer, and finally generate a fixed-length vector encoding. The above vector encodes various important information of the text description. This vector encoding can be the initial feature after encoding the text-form input data.
[0057] For another example, if the input data is an image, the CLIP model can process the input data using a convolutional neural network. For instance, using its image feature extractor, the input data is converted from an image to a vector representation. In the above process, the CLIP model can identify features such as visual elements, textures, colors, shapes, etc. in the input data, and synthesize the above features into a vector encoding that represents the content of the image. This vector encoding can be the initial feature after encoding the input data in the form of an image.
[0058] In the embodiment of the present invention, by using the CLIP model, it is ensured that the initial features after encoding the input data are comparable in the feature space, which is the key to realizing image search by text and image search by image. The pre-training mechanism of the CLIP model enables the CLIP model to have the ability to process a wide range of scenarios and concepts, while the pre-constructed feature database allows for quickly finding matching results during query, with a response time reaching the millisecond level. By encoding the input data through the CLIP model, the generated initial features represent the key features of the target object in the feature space, laying a solid foundation for subsequent accurate queries. The above process makes full use of the cross-modal encoding ability of the CLIP model and the pre-constructed feature database, effectively solving the problem of difficult target object recognition in open scenes and greatly improving the speed and accuracy of queries.
[0059] Step S106: Query for target features that match the initial features among the original image sample features corresponding to the original image samples in the database and the foreground image sample features corresponding to the foreground image samples.
[0060] In the technical solution provided in step S106 of the embodiment of the present invention, there is a mapping relationship between the original image samples and the foreground image samples. The original image samples are used to represent image samples with the image of an open scene sample as the background image sample and the image of an object sample as the foreground image sample, and the size of the object sample in the original image sample is less than a set threshold.
[0061] Optionally, the database can be referred to as a feature database (also known as a feature library) or a vector database. The original image sample can also be called the original file. The original image sample is the complete image stored in the database, representing the background in an empty scene and specific objects therein. The original image sample can be a complete image that has been processed by the Region of Interest (ROI) extraction module of the CLIP model and still contains information about the entire empty scene. The original image sample features can also be called the original data feature set. The foreground image sample can be an image sample containing object samples extracted from the original image sample by the ROI extraction module, and can also be called the ROI image. The foreground image sample mainly contains the image of the object sample in the original image sample, and the proportion of the object sample in the original image is less than a set threshold. Through the above selective extraction, it can be ensured that the CLIP model focuses on the key part of the image containing the object sample rather than the empty background, which is important for improving the search accuracy in specific scenarios (such as mine scenes).
[0062] Optionally, there is a mapping relationship between the original image sample and the foreground image sample, which means that each original image sample has a corresponding ROI image. The above mapping relationship is established in the database, providing a two-layer guarantee for the query process, that is, when querying, both the complete original image features and the ROI image features can be referred to, and the dual mechanism ensures the comprehensiveness and accuracy of image queries.
[0063] In this embodiment, the above set threshold is used to measure the critical value of the small proportion of the size of the object sample in the original image sample, and plays an important role in processing the original image sample and the foreground image sample. The set threshold directly affects the ability to identify the target object in the foreground information and the efficiency of extracting and highlighting these objects in an image with an empty scene as the background. The set threshold can consider the possible size change range of the target object in different scenarios. For example, in a mine environment, since objects ranging from individuals or tools to vehicles may all be search targets, the set threshold can be allowed to be adjusted within a certain range. For example, it can be set between 2% and 10% to adapt to different sizes of targets.
[0064] In addition, considering factors such as seasonal changes, weather conditions, time (such as the difference in light in the morning and evening), and the particularities of different mining areas, the set threshold can have a dynamic adjustment mechanism. For example, when the light is weak in the morning and at dusk, the set threshold can be appropriately lowered to capture more small targets; while under strong light conditions, the set threshold can be increased to avoid capturing too much noise. For certain specific categories of object samples, their proportion (ratio) in the original image samples may be naturally small or large. Therefore, the set threshold can be adjusted according to the different categories of object samples. For example, for targets of the "human" category, the set threshold can be relatively small, while for natural elements such as "water" or "sky" that occupy a large area, the set threshold can be set relatively large.
[0065] For example, in a mining scene, the size of the set object sample accounts for 5% of the original image sample. This means that objects with a proportion less than 5% are regarded as foreground information for ROI extraction and separate coding. For larger objects with an area proportion greater than or equal to 5%, the entire image can be directly used for coding, or additional detection can be performed on them to confirm whether they are the main search targets. When the input query text is "human", foreground image samples with an area proportion less than 5% and features such as shape and color matching human features are searched for. Thus, even in images with a vast background, human targets can be accurately captured without being interfered by a large amount of background information. When the input query is "large mining truck", considering that mining trucks usually occupy a relatively large proportion in the image, the threshold can be dynamically adjusted to a higher level, such as 15%. This can ensure that large targets such as trucks are correctly identified and coded and will not be overlooked even in images with multiple targets coexisting.
[0066] In the embodiment of the present invention, through the above-mentioned set threshold analysis and multiple sets of defined settings, the diverse search requirements in the mining scene can be effectively met, ensuring the accuracy and timeliness of the search results. At the same time, it can also be flexibly adjusted according to the changes in the actual scene and the specific needs of users, greatly enhancing the intelligence and practicality of image queries.
[0067] In this embodiment, after encoding the input data to obtain the initial features, target features matching the initial features can be queried from the features of the original image samples and the foreground image samples in the database.
[0068] Optionally, a query is performed based on the feature database to find target features matching the initial features encoded from the input data. The above process is the core of implementing the functions of "searching for images by image" and "searching for images by text". By using the features of the original image samples and the foreground image samples stored in the database, the similarity is calculated to locate target features (images or text descriptions) with a high correlation to the initial features.
[0069] Optionally, during the process of querying for target features that match the initial features, the initial features obtained by encoding the input data (text or image) can be compared with the original image sample features and foreground image sample features in the database.
[0070] For example, the target features that match the initial features (vectors with distances close to the vector of the initial features) can be determined from the original image sample features and foreground image sample features by calculating the distances between the vector of the initial features in the feature space and the vectors of the original image sample features and foreground image sample features in the feature space. For instance, the above distances can be represented by cosine similarity, Euclidean distance, etc.
[0071] Optionally, during the process of comparing the initial features with the original image sample features and foreground image sample features, the target features that match the initial features can be located from the database according to a preset similarity threshold or retrieval mechanism. The above target features represent images or text descriptions in the database with a similarity higher than the similarity threshold to the initial features.
[0072] Optionally, the constructed database not only stores the original image sample features of the original images but also includes the foreground image sample features of the ROI image samples, so that both background information and foreground information can be utilized simultaneously during the query process, which is particularly effective for target object recognition in open scenes.
[0073] In the embodiments of the present invention, by using a similarity search tool, even when dealing with large-scale data sets, the query process can achieve a response speed in milliseconds while ensuring the query accuracy. Combining ROI extraction can further improve the efficiency of the search task, especially in mining scenes where high-precision identification of specific objects or elements is required. By searching for the target features in the feature database that best match the initial features after encoding the input data, the functions of "searching for images by text" and "searching for images by image" are realized. The above process not only considers the complete information of the original image (background information + foreground information) but also particularly focuses on the object samples (foreground information), so that more accurate and faster search results can be provided when dealing with challenging open scenes. Through reasonable design and optimization, while improving the query efficiency, the high quality of the search results is also ensured, providing an effective solution for image and text retrieval in open scenes.
[0074] Step S108, label or output the target image corresponding to the target feature.
[0075] In the technical solution provided in step S108 of the embodiment of the present invention, the target image includes a target object. The target image can be a foreground image sample that conforms to the corresponding target feature, or an original image sample that includes foreground information and background information, or can also include a foreground image sample and an original image sample. The target image can be referred to as a search result.
[0076] In this embodiment, after querying the target feature that matches the initial feature from the original image sample features and foreground image sample features in the database, the target image corresponding to the target feature can be labeled or output.
[0077] Optionally, by querying the target feature that matches the initial feature after encoding the input data, the specific image corresponding to the above target feature, that is, the target image, is further determined. The target image can be a foreground image sample that only includes foreground information and conforms to the target feature, or an original image sample that includes foreground information and background information and conforms to the target feature, or a combination of the two. The above step plays a key linking role in the search process, converting the matching result in the feature space into an actual image for easy understanding and use by the user.
[0078] It should be noted that whether the above-generated target image is a foreground image sample that conforms to the target feature, or an original image sample, or a combination of the two, can be generated according to the user's needs to provide more personalized and accurate image search. If the user needs to focus on a specific target object or foreground details, a foreground image sample can be generated as the target image, only showing the part most relevant to the query, such as detected targets like humans, vehicles, etc., which helps the user quickly obtain information about the object of interest. If the user needs to understand the background environment where the target object is located, an original image sample will be generated to provide a more comprehensive view, including the target object and its surrounding environment. The user can also choose to view both foreground information and background information, and at this time, an image containing the above two types of information is generated.
[0079] Through the above process, it is possible to flexibly adapt to different scenarios and user needs, providing more accurate search results that meet user expectations. For example, in a mine scenario, if the user wants to find a specific type of machinery or worker, an image containing only foreground information can be requested; while if the user needs to understand the environment where the machinery operates or the specific location where the worker works, the original image sample will be more useful. The above flexibility improves the search efficiency and user satisfaction. In summary, the technical process of generating the target image according to the user's needs not only solves the problems of low accuracy and poor practicability of the search results, but also realizes a more efficient and intelligent information retrieval function through feature space matching and flexible control of image generation.
[0080] Optionally, after the target feature is located, the target image can be annotated. For example, the found target object, the region where the target object is located, and the category information of the target object (such as person, vehicle, accumulated water, etc.) can be marked on the image. For the "image search by image" task, annotation can further confirm the accuracy of the search results, while for the "image search by text" task, annotation helps identify the key parts in the image. The annotated target image can be used for subsequent model training as positive samples, enabling the CLIP model to better understand the relationship between images and texts in specific scenarios.
[0081] Optionally, after the annotation is completed or when annotation is not required, the target image can be output. The output target image can be single or multiple, depending on the search strategy and user requirements. The output target image provides the visual or text information required by the user, meeting the basic needs of "image search by image" and "image search by text".
[0082] Optionally, in the search results, the target image can include a foreground image and a background image, reflecting the flexibility and comprehensiveness of the search method when dealing with mine scenes. By combining foreground and background information, users can more comprehensively understand the context of the search results. Especially in cold start tasks, the above combination can provide richer training data, helping the model quickly adapt to new scenarios. The output target image should highly match the input data to ensure the accuracy and reliability of the search. For example, the search method in the embodiments of the present invention can improve the matching accuracy and increase the robustness to complex scenarios by combining ROI detection, bad image detection, and CLIP encoding. When dealing with abnormal images such as flower screens and green screens, the accuracy of the method is particularly prominent, ensuring that the quality of the output results is not affected.
[0083] In the embodiments of the present invention, the search process loop is completed by annotating or directly outputting the target image that matches the target feature. The above process not only verifies the accuracy of the search results but also provides data support for the model optimization used in the subsequent embodiments of the present invention, which is crucial for improving the efficiency and accuracy of image search in mine scenes. By comprehensively considering foreground and background information and ensuring the quality and relevance of the output image, accurate search for target objects in specific scenarios is achieved.
[0084] In the embodiments of the present invention, the above target image can be used to train a perception model. Among them, the trained perception model can be used to identify the perception data collected by the driverless vehicle to obtain perception results.
[0085] Optionally, the perception model can be a deep learning model. For example, it can be an architecture such as a convolutional neural network, which is used to process the perception data provided by the sensors of the autonomous vehicle. The perception data can include objects in an empty scene, which can be images or texts of the objects in the empty scene. For example, camera images, etc. The perception model can convert the above perception data into a high-level understanding of the environment to obtain corresponding perception results. That is to say, the perception result is the recognition and understanding of the environment around the autonomous vehicle by the perception model. For example, if the perception data is a camera image, the corresponding perception result can be pedestrians, vehicles, obstacles, etc. recognized in the camera image. More specifically, the perception result can even be information such as the category label, position coordinates, and size of the objects in the camera image. There is no specific limitation here. It should be noted that the perception result can correspond to the perception data. Through the generation of the perception result, it can be shown that the perception model can accurately recognize various information that matches the perception data.
[0086] Optionally, the autonomous vehicle (driverless vehicle) can be the carrier for applying the trained perception model, responsible for collecting perception data in the real world and making driving decisions based on the recognition results of the model. The autonomous vehicle is equipped with a variety of sensors, including but not limited to cameras, lidar, millimeter-wave radars, ultrasonic sensors, etc. The above sensors continuously collect perception data of the surrounding environment and provide inputs for the perception model. Based on the perception results obtained by the perception model for recognizing the perception data, the autonomous vehicle can make decisions, such as adjusting speed, changing lanes, avoiding obstacles, etc., and then realize autonomous driving through the actuators (such as motors, steering systems) on the autonomous vehicle.
[0087] In this embodiment, the above target image can be used to train the perception model. And the trained perception model is used to recognize the perception data collected by the autonomous vehicle to obtain the final perception result.
[0088] Optionally, this embodiment is a key link in building the autonomous vehicle perception system, involving the training and application of the perception model to accurately recognize the perception data collected by the autonomous vehicle and generate perception results, thereby providing a basis for the vehicle's decision-making and navigation.
[0089] Optionally, the perception model iteratively processes the target image repeatedly, attempts to predict the correct object category, position, and other attributes from the input target image, and at the same time compares with the labeled data through a loss function, continuously adjusting the model parameters to reduce the prediction error until the model reaches a satisfactory recognition accuracy. The perception model can be various deep learning architectures, such as convolutional neural networks, recurrent neural networks, etc.
[0090] Optionally, during the operation of the driverless vehicle, the perception data can be continuously fed into the perception model. The perception model performs real-time recognition and classification on the acquired perception data according to the patterns learned during training, and outputs a description of the vehicle's surrounding environment, including the position, type, size of the objects, and their dynamic changes (such as the movement direction and speed).
[0091] Optionally, the target image obtained by the above method can be used to train the perception model, and the trained perception model can be used to identify the perception data collected by the driverless vehicle, so as to obtain the perception result. In this embodiment, since the target image obtained by the above method can greatly reduce the interference of irrelevant background information, using this target image to train the perception model can also enable the perception model to avoid the interference of irrelevant background information and focus on identifying and analyzing the key foreground information in the target image. The perception model trained by the above method can also avoid the interference of irrelevant information in the perception data and focus on important information when applied to the scenario of perception data recognition, so as to obtain the perception result corresponding to the important information. Furthermore, it ensures the perception effect of the driverless vehicle on the surrounding environment in an open scenario. It solves the technical problem of low perception accuracy in driverless driving and achieves the technical effect of improving the perception accuracy of driverless driving.
[0092] In the technical solutions provided in steps S102 to S108 of the embodiment of the present invention, the input data of the target object in the to-be-query open scenario can be obtained, and the input data is encoded to obtain the initial features. The target features matching the above initial features can be queried from the original image sample features and foreground image sample features in the database. And the target image including the target features and the target object can be output or labeled. In this embodiment, in the face of the challenge that the target object accounts for a small proportion in the open scenario, the foreground image samples having a mapping relationship with the original image samples in the open scenario can be pre-identified and extracted, and a database is established by combining the foreground image sample features and the original image sample features. In the query stage, through the efficient comparison of the initial features obtained by encoding the input data with the original image sample features and foreground image sample features in the database, even in the cold start scenario lacking specific samples, the image matching the to-be-query target object can be accurately identified, significantly improving the accuracy and efficiency of image query in the open scenario. Through the above feature encoding of the input data and database query strategy, the interference of irrelevant background information is greatly reduced, and the sensitivity to rare targets is improved, thus ensuring the image query effect in the open scenario. It solves the technical problem of low query accuracy of images and achieves the technical effect of improving the query accuracy of images.
[0093] The embodiments of the present invention will be described in detail below in combination with the above steps.
[0094] As an alternative embodiment, in step S106, querying for a target feature that matches the initial feature among the original image sample features corresponding to the original image samples in the database and the foreground image sample features corresponding to the foreground image samples includes: querying for original image sample features in the database whose similarity to the initial feature is greater than the similarity threshold, and / or querying for foreground image sample features in the database whose similarity to the initial feature is greater than the similarity threshold; and determining the queried original image sample features whose similarity to the initial feature is greater than the similarity threshold and / or the queried foreground image sample features as the target features that match the initial feature.
[0095] In this embodiment, during the process of querying for a target feature that matches the initial feature among the original image sample features and the foreground image sample features in the database, original image sample features whose similarity to the initial feature is greater than the similarity threshold can be queried in the original image sample features. Alternatively, foreground image sample features whose similarity to the initial feature is greater than the similarity threshold can be queried in the foreground image sample features. The queried foreground image sample features and / or original image sample features whose similarity to the initial feature is greater than the similarity threshold can be determined as the target features. Herein, the similarity threshold is an important parameter for controlling the strictness of the matching. A lower threshold may result in more incompletely relevant features being regarded as matches, thereby increasing the number of target images obtained by the query, but may reduce the relevance between the target images and the initial feature. A higher similarity threshold, on the contrary, may narrow the range of target images, improve the accuracy of the target images, but may also cause some features with high relevance to be excluded due to minor differences. Therefore, reasonably setting the similarity threshold is crucial for ensuring the accuracy and practicality of the query results.
[0096] Optionally, this embodiment refines the feature matching process, introduces the concept of the similarity threshold, and more precisely determines the target features that match the initial feature after encoding the input data. The image query process of the embodiment of the present invention is divided into two parts: one is to perform matching in the original image sample features, and the other is to perform matching in the foreground image sample features. By setting the similarity threshold, sample features (foreground image sample features and / or background image sample features) whose similarity to the initial feature is high enough can be filtered out, and the above sample features can be determined as the target features that match the initial feature.
[0097] Optionally, in the original image sample features, the query process involves calculating the similarity between each original image sample feature in the database and the initial feature after encoding the input data. If the similarity between a certain original image sample feature and the initial feature is greater than a preset similarity threshold, then this feature is regarded as a matching feature, that is, a part of the target feature. The above process ensures that in the case of a complex background or an empty scene, the target feature highly relevant to the initial feature can still be found.
[0098] Optionally, the query process for the foreground image sample features is similar to that of the original image sample features, but the object is the foreground image sample features obtained after being processed by the ROI extraction module. Since the above initial features are concentrated on the target object, a more accurate matching result can be provided in the similarity calculation. If the similarity between a certain foreground image sample feature and the initial feature is greater than the similarity threshold, similarly, the above foreground image sample feature can be determined as the target feature.
[0099] Optionally, in the comprehensive query process, matching is performed not only in the original image sample features but also in the foreground image sample features, increasing the dimension and possibility of matching. The query result can include any one or both of the two features. As long as the similarity between the above features and the initial feature is greater than the similarity threshold, it can be determined as the target feature that matches the initial feature.
[0100] Optionally, by setting similarity thresholds in the original image sample features and the foreground image sample features respectively, the embodiments of the present invention can more accurately locate the target features highly relevant to the input data when processing complex mine scenes. The implementation effects of the above method can be reflected in the following aspects: Using the similarity threshold for filtering can significantly reduce the number of features that need to be further processed, thereby improving the search speed and achieving millisecond-level response. By concentrating on the results with a similarity exceeding the threshold with the input features, the search method can reduce mis-matching and improve the accuracy of the results. Especially when processing scenes with an empty background and few targets, the foreground image sample features extracted by ROI can provide more accurate matching. It allows users to adjust the similarity threshold according to actual needs to obtain query results that meet specific criteria. At the same time, the target images corresponding to the queried target features can be directly used for model training or in mine scenes, increasing the practicality of the search method.
[0101] In the embodiments of the present invention, by setting similarity thresholds in the original image sample features and the foreground image sample features of the database, accurate matching of the target features is achieved. The above process not only improves the efficiency and accuracy of the query but also provides flexible parameter adjustment to adapt to different search requirements. By comprehensively considering background and foreground information, it can effectively target complex mine scenes and provide high-quality image and text search results.
[0102] As an alternative embodiment, the method further includes: extracting the foreground from the original image sample to obtain a foreground image sample; establishing a mapping relationship between the foreground image sample and the original image sample; and storing the foreground image sample features and the original image sample features in a database according to the mapping relationship.
[0103] In this embodiment, the foreground can be extracted from the original image sample to obtain a foreground image sample. A mapping relationship can be established between the foreground image sample and the original image sample, so that the foreground image sample features and the original image sample features can be stored in a database according to the mapping relationship.
[0104] Optionally, this embodiment refines the process of data preprocessing and database construction, focusing on the extraction of foreground image samples and the establishment of the mapping relationship with the original image samples, aiming to enhance the accuracy and efficiency of the search task.
[0105] Optionally, the foreground is extracted from the original image sample to obtain a foreground image sample. The above process utilizes an ROI extraction module, based on deep learning, such as the You Only Look Once Version 11 (abbreviated as YOLO11) series or other object detection models, to automatically identify and separate the objects or regions of interest in the image. For a mine scene, this method can effectively focus on key elements such as trucks and personnel, while ignoring large areas of empty background regions, thereby improving the pertinence and effectiveness of feature encoding.
[0106] Optionally, a mapping relationship is established between the foreground image sample and the original image sample. The above steps ensure that each foreground image sample can be associated with the corresponding original image sample, providing an important basis for subsequent search operations. For example, when the foreground image sample features that best match the input features are found, the original image sample can be traced back through the mapping relationship to obtain the complete scene information.
[0107] Optionally, the foreground image sample features and the original image sample features are stored in a database according to the mapping relationship. This step is crucial for constructing a feature database (or vector database). By storing the feature information at two levels, namely the original image sample features of the complete scene and the foreground image sample features of the focused objects, rich data resources are provided for subsequent feature comparison and query operations. Using efficient data management and search tools, it can be ensured that the storage and retrieval processes are both fast and accurate.
[0108] In the embodiments of the present invention, by storing the features of the original image and the foreground image simultaneously, multi-scale encoding and retrieval of image information can be achieved, taking into account both the global background and local details, which is particularly important for search tasks in complex scenarios. Establishing the mapping relationship between the foreground image samples and the original image samples enables the system to flexibly switch perspectives, starting either from the local (foreground image) or returning to the global (original image). This dual perspective provides more possibilities when dealing with specific requirements. Storing the features of the foreground and original images according to the mapping relationship optimizes the database structure, making the retrieval process more efficient. Especially for scenarios with an empty background such as mine scenes, by preferentially retrieving the features of the foreground image, the key objects can be located more quickly, thus improving the response speed and accuracy of the search.
[0109] In summary, by introducing the extraction of foreground image samples and the establishment of the mapping relationship with the original image samples, not only the content richness of the database is enhanced, but also the data storage and retrieval mechanisms are optimized. The above method significantly improves the accuracy and efficiency of the search task when processing complex mine scene images, demonstrating the practicality of the embodiments of the present invention in specific industry applications. Through careful data preprocessing and database construction, the embodiments of the present invention successfully solve the key problem of how to perform efficient and accurate image and text search in the cold start situation without historical data or basic model support.
[0110] As an optional embodiment, foreground extraction is performed on the original image sample to obtain the foreground image sample, including: performing foreground detection on the original image sample to obtain detection information; and extracting the foreground image sample from the original image sample based on the detection information.
[0111] In this embodiment, during the process of performing foreground extraction on the original image sample to obtain the foreground image sample, foreground detection can be performed on the original image sample to obtain the detection result. The foreground image sample can be extracted from the original image sample based on the detection result.
[0112] Optionally, the extraction process of the foreground image sample is divided into two main steps: foreground detection and foreground image extraction based on the detection result. The above process aims to isolate the region of interest (ROI) from the original image to provide more focused and refined image features, which is particularly suitable for scenarios with an empty background such as mine environments, thereby improving the accuracy of feature encoding and subsequent search tasks.
[0113] As an alternative embodiment, foreground detection is performed on the original image sample to obtain detection information, including: invoking an information extraction model to perform foreground detection on the original image sample to obtain detection information, where the information extraction model is trained using detection information samples of object samples in an empty scene; based on the detection information, extracting a foreground image sample from the original image sample, including: using the information extraction model to analyze the detection information and the original image sample to obtain the foreground image sample.
[0114] In this embodiment, in the process of performing foreground detection on the original image sample to obtain detection information, an information extraction model can be invoked to perform foreground detection on the original image sample to obtain detection information. The information extraction model can be used to analyze the detection information and the original image sample to obtain the foreground image sample. Among them, the information extraction model can be trained using detection information samples of object samples in an empty scene and can be an ROI extraction module.
[0115] Optionally, the foreground detection of the original image sample and the creation of the foreground image sample are refined into two key steps: using the information extraction model to perform foreground detection to obtain detection information; based on the above detection information, using the model to analyze the original image and accurately extract the foreground image sample from it. The above process particularly emphasizes the training and use of the information extraction model to meet the object detection requirements in an empty scene.
[0116] Optionally, the information extraction model should be optimized for object detection in an empty scene. This means that when the model is trained, object samples and detection information samples specific to this type of scene are used. For example, in a mine scene, the model may be well-trained to identify key objects such as trucks, excavation equipment, and personnel, even if they occupy a small proportion in the vast space.
[0117] Optionally, in practical applications, by invoking the trained information extraction model, the original image sample is analyzed to identify the foreground objects therein. The information extraction model can output detection information such as the position (bounding box), category, and confidence of each identified object, and the above detection information is used to guide the subsequent ROI extraction process.
[0118] Optionally, using the obtained detection information, the region of interest (ROI) in the image is located. For example, the bounding box of each object is determined and used as the basis for cropping the foreground image sample. From the original image sample, according to the position information of each ROI, the region containing the object is accurately cropped to form the foreground image sample. The above foreground image sample focuses on key objects and removes a large amount of background information, making it suitable for the next step of feature encoding and storage.
[0119] In the embodiments of the present invention, the information extraction model is trained based on object samples in an empty scene, which means that it has higher detection accuracy and robustness when processing environments with relatively less background information such as mine scenes. The model can effectively identify key objects in the foreground, even if the objects occupy a small proportion in the image. By performing foreground detection and ROI extraction on the original image, this embodiment can significantly reduce the amount of data during feature encoding, thereby improving the efficiency of the search task. At the same time, due to focusing on the key areas, the results of feature encoding are more accurate, which helps to improve the relevance and accuracy of the search results. Only performing feature encoding on the foreground image samples avoids processing the entire image dataset, thus reducing the consumption of computing resources, especially when dealing with large-scale image sets, which is particularly important.
[0120] In summary, by using the information extraction model optimized for empty scenes to perform foreground detection and accurately extracting foreground image samples based on the detection information, the problem of efficient and accurate image search in mine scenes with an empty background is effectively solved. The above method not only improves the accuracy of ROI detection and extraction, but also optimizes the efficiency of subsequent feature encoding and search tasks, demonstrating its significant advantages in specific industry applications. By focusing on key information, it can provide more accurate search results, meeting the requirements for rapid and accurate retrieval of image information in mine scenes.
[0121] As an optional embodiment, according to the mapping relationship, storing the foreground image sample features and the original image sample features in a database includes: respectively performing feature encoding on the foreground image sample and the original image sample with a mapping relationship in the target feature space to obtain the foreground image sample features and the original image sample features with a mapping relationship; storing the foreground image sample features and the original image sample features with a mapping relationship in the database.
[0122] In this embodiment, during the process of storing the foreground image sample features and the original image sample features in the database according to the mapping relationship, the foreground image sample and the original image sample with a mapping relationship can be respectively subjected to feature encoding in the target feature space, so as to obtain the foreground image sample features and the original image sample features with a mapping relationship. And store the foreground image sample features and the original image sample features with a mapping relationship in the database.
[0123] Optionally, perform feature encoding on the foreground image sample and the original image sample with a mapping relationship, and then store the encoded features in the database. The above process ensures that the key information of the image is stored in the database, while maintaining the relevance between the original image and its foreground, providing an efficient and accurate basis for subsequent queries.
[0124] Optionally, a target feature space is defined, which is a vector space that maps image and text information into a unified dimension and format, facilitating efficient similarity calculation and search. In the embodiments of the present invention, the CLIP model plays a key role and can transform image and text information into vectors in this target feature space.
[0125] Optionally, for each foreground image sample, use the CLIP model for feature encoding to obtain its representation in the target feature space. The feature encoding centrally reflects the key object information in the foreground, and through the ROI extraction process, these samples usually contain the main entities in the scene, reducing the interference of background noise. Similarly, perform feature encoding on the original image sample to obtain its representation in the target feature space. The above process preserves the global information of the image and can provide a complete view of the scene even when the background is empty.
[0126] Optionally, utilize the mapping relationship between the foreground image sample and the original image sample to store the encoded foreground image sample features and the original image sample features in the database. The above mapping relationship ensures that each feature record in the database can be traced back to the corresponding original image and foreground image, providing a two-way link for search and query. Storing the encoded features in the database, the above process involves vector database technologies such as Facebook AI Similarity Search (abbreviated as FAISS), which can efficiently manage and search high-dimensional vector data. When storing, each feature vector will be accompanied by the number or other identification information of the image it maps to, to maintain the association between the image features and the actual images.
[0127] Optionally, through feature encoding by the CLIP model, it can ensure the representation of image and text information in a unified vector space, making it possible to query images based on text or query text based on images, expanding the retrieval ability of the database. Saving the features with mapping relationships in the database means that during retrieval, not only the features themselves can be considered, but also the association between the features can be taken into account, which provides additional flexibility and accuracy for processing complex scenarios such as image retrieval in a mine environment. Using vector database technology for storage can achieve fast retrieval on a large-scale dataset, and even query the most relevant image within milliseconds, which greatly improves the system response speed and user experience.
[0128] In the embodiments of the present invention, through the above method, not only can the foreground image samples and the original image samples with mapping relationships be efficiently feature-encoded, but also the encoded features can be stored in the database, maintaining the integrity and relevance of the image information. The above process not only optimizes the data storage structure, improves the retrieval efficiency, but also ensures that when dealing with scenarios such as mine scenes where the background is empty and there are few targets, accurate and fast search results can be provided, reflecting its advantages and value in specific industry applications. By combining ROI extraction, feature encoding, and vector database technology, the embodiments of the present invention achieve effective management and retrieval of image information in mine scenes, providing strong technical support for related requirements.
[0129] As an optional embodiment, step S104, encoding the input data to obtain initial features, includes: performing feature encoding on the input data in the target feature space to obtain initial features.
[0130] In this embodiment, during the process of encoding the input data to obtain initial features, the input data can be feature-encoded in the target feature space to obtain initial features.
[0131] Optionally, the goal of feature-encoding the input data is to convert the input data into a format that can be operated and compared in a specific target feature space. The target feature space is a predefined and unified multi-dimensional vector space, in which different types of input data (such as text, images, etc.) can be represented as vector points, so that cross-modal data (for example, images and text) can be analyzed or matched for similarity under the same conditions.
[0132] Optionally, the selection and definition of the target feature space are based on the design of the model and the requirements of the application scenario. For example, the CLIP model is an example designed to map images and text into the same vector space. Its target feature space is universal and can accommodate high-dimensional representations of various images and texts, facilitating subsequent retrieval and similarity calculation.
[0133] Optionally, whether the input data is an image or text, it can be feature-encoded through a deep learning model to obtain its representation in this target feature space. For example, if the input data is an image, then the image encoder in the CLIP model can process the above input data, extract key visual features, and convert them into a fixed-length vector, which is called the initial feature of the image. If the input is text, then the text encoder will be responsible for converting the semantic information into vector form to form the initial feature of the text.
[0134] Optionally, by encoding data of different modalities as vectors in the target feature space, the consistency and comparability of image and text data are ensured. This means that whether the input is an image or text, it can be processed in the same way, thus simplifying the subsequent retrieval process. The process of feature encoding is also a process of data compression and feature extraction. The model automatically learns the key features in the input data through deep learning techniques and transforms them into a vector, enabling the retrieval process to be faster and more efficient. Once the input data is encoded as a vector in the target feature space, the same retrieval algorithm and strategy can be used for searching. Whether it is to query the matching image corresponding to the text or the relevant text description corresponding to the image, it is completed on the same platform, increasing the flexibility and generality of the system.
[0135] In the embodiment of the present invention, through the above feature encoding process, the input data is transformed into initial features in the form of vectors in the target feature space. The above transformation not only ensures the cross-modal consistency of the data, but also improves the efficiency and accuracy of data retrieval through feature compression and representation. Finally, whether the input is an image or text, similarity matching and searching can be performed on a unified retrieval platform, providing key technical support for realizing efficient and accurate image search by text and image search by image services. The above process makes full use of the capabilities of modern deep learning models, especially cross-modal learning models such as CLIP, which can map images and text to the same feature space, greatly improving the performance and user experience of image information retrieval in mine scenarios.
[0136] As an optional embodiment, the method further includes: performing quality detection on the original image sample to obtain the quality detection result of the original image sample; performing foreground extraction on the original image sample to obtain the foreground image sample, including: in response to the quality detection result indicating that the quality of the original image sample is qualified, performing foreground extraction on the original image sample to obtain the foreground image sample.
[0137] In this embodiment, the quality of the original image sample can be detected to obtain the quality detection result of the original image sample. If the quality detection result of the original image sample is qualified, the foreground of the original image sample can be extracted to obtain the foreground image sample.
[0138] Optionally, the quality of the original image sample is detected to ensure that subsequent image analysis and feature extraction can be performed based on high-quality image data. Only when the image quality detection result indicates that the image quality is qualified, will the foreground of the above high-quality original image sample be further extracted to focus on the key objects in the image, thereby improving the effectiveness of feature encoding and the accuracy of the search task. The above process aims to optimize the image dataset, eliminate low-quality images, and improve the overall performance of image queries.
[0139] Optionally, quality detection can be performed on the original image samples. A dedicated Image Quality Assessment (IQA) model or algorithm can be used. The above-mentioned image quality assessment model can effectively identify key quality indicators such as image sharpness, contrast, color saturation, and noise level to determine whether the image is suitable for subsequent analysis and encoding.
[0140] Optionally, by analyzing the output of the quality detection model, the quality detection results of each image can be obtained. If the image quality meets a certain standard or threshold, the result will be marked as qualified in quality; otherwise, it will be marked as unqualified in quality.
[0141] Optionally, after obtaining the quality detection results of the original image samples, check whether the quality detection results indicate that the quality of the original image samples is qualified. The above judgment step is a key control point in the process and determines whether to perform foreground extraction on the image.
[0142] Optionally, for images with qualified quality detection results, the foreground extraction module can be called, and an object detection model such as YOLO11 can be used to identify and locate foreground objects. The above-mentioned foreground extraction model can accurately detect key objects in the image and define their bounding boxes, thereby realizing the extraction of the foreground region. After the foreground region is identified and located, the above region is cropped from the original image sample to form a foreground image sample. The above foreground image sample contains the main body or important objects of the image, removing a large amount of background information and providing more focused data for subsequent feature encoding.
[0143] In the embodiments of the present invention, through pre-quality detection, low-quality images can be effectively filtered out, avoiding feature extraction errors and search result deviations caused by poor image quality itself. On the basis of ensuring qualified image quality, using an advanced object detection model to analyze the image can more accurately locate and extract the foreground region in the image, reducing the interference of the background on feature encoding. Since only images with qualified quality are subjected to foreground extraction and feature encoding, the efficiency of feature encoding is further improved, and at the same time, because it focuses on the key parts of the image, the accuracy of the encoding result is also improved.
[0144] In summary, by implementing image quality detection and foreground extraction in response to qualified quality, the embodiments of the present invention can, when processing the original image samples, eliminate low-quality images one step ahead, ensuring that subsequent image analysis and feature encoding work are carried out on the basis of high-quality data. This not only improves the accuracy of feature encoding and the efficiency of search tasks, but also enables the system to focus more on the key objects in the image, reduces the interference of background noise, and has a significant optimization effect on image information retrieval and processing in the mine scene. The above process design reflects the strict control of the quality of input data and the accuracy of extracting key image information, providing strong support for realizing efficient and accurate text-to-image and image-to-image search services.
[0145] As an alternative embodiment, the quality detection result includes a first quality detection result. Performing quality detection on the original image sample to obtain the quality detection result of the original image sample includes: respectively determining the color information of the original image sample in different color spaces to obtain multiple color information; in response to each piece of color information among the multiple color information being within the normal color threshold range, determining that the first quality detection result is that the quality of the original image sample is qualified; in response to at least one of the multiple color information not being within the normal color threshold range, determining that the first quality detection result is that the quality of the original image sample is unqualified.
[0146] In this embodiment, in the process of performing quality detection on the original image sample to obtain the quality detection result of the original image sample, the color information of the original image sample in different color spaces can be respectively determined to obtain multiple color information. If each piece of color information among the multiple color information is within the normal color threshold range, it can be determined that the first quality detection result is that the quality of the original image sample is qualified. If at least one of the multiple color information is not within the normal color threshold range, then it can be determined that the first quality detection result is that the quality of the original image sample is unqualified. The quality detection result can include the first quality detection result. The color space can include the Blue Green Red (BGR) channel, the Grayscale (GRAY) space, and the L*a*b (LAB) space.
[0147] Optionally, this embodiment proposes a comprehensive image quality detection framework. This framework is based on the statistical characteristics of multiple color spaces (BGR, GRAY, LAB) to evaluate whether an image is suitable for subsequent feature encoding and image retrieval tasks. The above method can capture the inherent properties of an image under different color expressions, thereby more accurately determining whether its quality meets the requirements.
[0148] Optionally, analyze the color information of the original image sample in the BGR color space, obtain the values of the blue, green, and red color channels for each pixel, and calculate the average value and standard deviation of the entire image on each color channel. BGR channel analysis can reveal the uniformity and richness of the image color distribution.
[0149] Optionally, convert the original image sample to the grayscale space, that is, only retain one channel representing brightness. In the GRAY space, calculate the average grayscale value and the standard deviation of the grayscale values, which helps to evaluate the contrast and detail visibility of the image.
[0150] Optionally, analyze the original image sample in the LAB color space, which divides colors into brightness (L) and two color channels (a, b). By calculating the average values and standard deviations of the L, a, and b channels, an in-depth understanding of the image's brightness, color saturation, and balance can be obtained.
[0151] Optionally, after obtaining the statistical values in the above three color spaces, a series of normal color thresholds can be set, and the above normal color thresholds are used to determine whether the color information of the image in each color space is within the normal range.
[0152] Optionally, for each image, if the statistical values (such as the average value and standard deviation) of its color information in the analyzed color spaces (BGR, GRAY, LAB) all fall within the set normal color threshold range, it is determined that the image passes the quality inspection, that is, the first quality inspection result is qualified. On the contrary, if the color information in at least one color space exceeds the normal threshold range, it is determined that the image quality is unqualified.
[0153] In the embodiment of the present invention, a comprehensive quality inspection mechanism is constructed by analyzing the color information of the image in multiple color spaces. The above method can evaluate the image from multiple dimensions such as brightness, color saturation, and color balance, ensuring that the image is clear, rich, and balanced visually, and is suitable for feature extraction and image retrieval tasks. The setting of the normal color threshold needs to consider the requirements of the specific application scenario. For example, images in a mine scene may have specific color ranges and brightness requirements. By adjusting the normal color threshold, it can be adaptively determined whether the image quality meets the requirements of the current task, enhancing the flexibility and practicality of the method. By making a comprehensive judgment in multiple color spaces, the probability of false detection in a single color space can be effectively reduced, improving the accuracy and stability of the entire quality inspection process. Even for images under complex lighting conditions or with specific color backgrounds, a relatively objective quality evaluation can be obtained.
[0154] In summary, through the image quality detection method based on multiple color spaces, a comprehensive color characteristic analysis of the original image sample can be carried out, and the quality status of the image can be accurately determined according to the set normal color threshold. This method can not only ensure that the image quality for feature encoding and image retrieval meets the standards, but also adapt to different application scenarios, improving the efficiency of image processing and the reliability of the final retrieval results. In the mine scene and other image processing and retrieval tasks with complex environments, the above quality detection method is particularly crucial because it can exclude images that are not suitable for analysis due to color distortion or insufficient contrast, thus enhancing the performance of the entire image retrieval system and the user experience.
[0155] As an alternative embodiment, the quality detection result includes a second quality detection result. When performing quality detection on the original image sample to obtain the quality detection result of the original image sample, it includes: determining a first confidence level of the original image sample, where the first confidence level is used to represent the degree of possibility that the quality of the original image sample is qualified; in response to the first confidence level being greater than the first confidence level threshold, determining that the second quality detection result is that the quality of the original image sample is qualified; in response to the first confidence level being less than or equal to the first confidence level threshold, determining that the second quality detection result is that the quality of the original image sample is unqualified; or, determining a second confidence level of the original image sample, where the second confidence level is used to represent the degree of possibility that the quality of the original image sample is unqualified; in response to the second confidence level being greater than the second confidence level threshold, determining that the second quality detection result is that the quality of the original image sample is unqualified; in response to the second confidence level being less than or equal to the second confidence level threshold, determining that the second quality detection result is that the quality of the original image sample is qualified.
[0156] In this embodiment, during the process of performing quality detection on the original image sample to obtain the quality detection result of the original image sample, the first confidence level of the original image sample can be determined. If the first confidence level is greater than the first confidence level threshold, it can be determined that the second quality detection result is that the quality of the original image sample is qualified. If the first confidence level is less than or equal to the first confidence level threshold, it can be determined that the second quality detection result is that the quality of the original image sample is unqualified. It is also possible to determine the second confidence level of the original image sample. If the second confidence level is greater than the second confidence level threshold, then it can be determined that the second quality detection result is that the quality of the original image sample is unqualified. If the second confidence level is less than or equal to the second confidence level threshold, it can be determined that the second quality detection result is that the quality of the original image sample is qualified. Among them, the quality detection result includes the second quality detection result. The first confidence level can be the good image confidence level of the original image sample. The second confidence level can be the bad image (problem image) confidence level of the original image sample. The first confidence level threshold can be represented by representation. The second confidence level threshold can be represented by representation.
[0157] Optionally, this embodiment proposes a method for evaluating image quality based on the prediction confidence of a deep learning model. By calculating the confidence of a good image or a bad image for the original image sample and comparing it with preset thresholds (the first confidence threshold and the second confidence threshold), it is determined whether the image quality is qualified. The above method combines the prediction ability of modern machine learning models and statistical decision theory, providing a flexible and efficient way for image quality detection.
[0158] Optionally, a trained deep learning model, such as MobileNetV2 or other networks, is used to predict the likelihood that the original image sample is a good image. The higher the first confidence, the more likely the image quality is determined to be qualified, that is, the image is clear, without serious distortion, has good color balance, etc. The same deep learning model is used, but this time the prediction target of the above model is to evaluate the confidence that the image is a bad image. The higher the second confidence, the more likely the image quality is determined to be unqualified, that is, the image is blurred, has serious color distortion, has problems such as a mosaic screen or a green screen.
[0159] Optionally, for the original image sample, if the first confidence (good image confidence) is greater than the set first confidence threshold, then it can be determined that the image has passed the quality inspection, that is, the second quality inspection result is qualified. On the contrary, if the first confidence is less than or equal to the first confidence threshold, it indicates that the image quality is poor, and the second quality inspection result is unqualified. If the second confidence (bad image confidence) is used for determination, when the second confidence is greater than the set second confidence threshold, the image quality is also determined to be unqualified. If the second confidence is less than or equal to the second confidence threshold, the image quality is determined to be qualified.
[0160] Optionally, the training of the deep learning model is based on a large number of labeled image datasets. The images in the above datasets have been classified into two categories: good images and bad images according to their quality. Through training, the deep learning model can learn the key features to distinguish good images from bad images, so as to accurately predict the quality status of new images in the test stage. The setting of the first confidence threshold and the second confidence threshold needs to be based on the performance during model training and the requirements of the actual application scenario. The above two thresholds can be determined through cross-validation and performance evaluation to balance the false detection rate and the false alarm rate, ensuring the accuracy and stability of quality detection. It can also be combined with traditional algorithms. When the output of the deep learning model is inconsistent with traditional methods (such as statistical methods based on color space), the system will be based on the set confidence threshold and carry out cross-validation to further improve the detection accuracy.
[0161] In the embodiments of the present invention, through the above-mentioned image quality detection method based on confidence, the prediction ability of the deep learning model can be utilized, combined with statistical decision theory, to effectively evaluate whether the quality of the original image sample is qualified. The above method can not only adapt to various image quality and defect types, but also has a high degree of automation and flexibility, and can significantly improve the efficiency and accuracy of image quality detection. In the mine scenario, the above method is particularly important because it can ensure that the images used for feature encoding and image retrieval have sufficient clarity and information content, thereby improving the overall performance and reliability of the image retrieval system. At the same time, by setting a reasonable confidence threshold, the detection results can be further optimized, misjudgments can be reduced, and high-quality images can be correctly identified and used.
[0162] As an alternative embodiment, in the case where the first quality detection result is different from the second quality detection result, the method further includes: in response to the first confidence being greater than the first confidence threshold, outputting the second quality detection result; in response to the first confidence being less than or equal to the first confidence threshold, outputting the first quality detection result; or in response to the second confidence being greater than the second confidence threshold, outputting the second quality detection result; in response to the second confidence being less than or equal to the second confidence threshold, outputting the first quality detection result.
[0163] In this embodiment, in the case where the first quality detection result is different from the second quality detection result, if the first confidence is greater than the first confidence threshold, the second quality detection result can be output. If the first confidence is less than or equal to the first confidence threshold, the first quality detection result can be output. If the second confidence is greater than the second confidence threshold, the second quality detection result can be output. If the second confidence is less than or equal to the second confidence threshold, the first quality detection result can be output.
[0164] Optionally, this embodiment proposes a decision-making process based on the fusion of two different quality detection methods. This process independently applies two quality detection methods, the first quality detection based on statistical thresholds and the second quality detection based on deep learning confidence prediction. Subsequently, according to the comparison result of the confidence and the threshold, it is decided which quality detection result to finally adopt as the basis for judging the image quality. The above method can combine the advantages of traditional statistical methods and deep learning models, and provide a more robust and accurate decision-making mechanism for image quality detection.
[0165] Optionally, the steps of the first quality inspection adopt traditional image analysis methods, such as calculating color statistical information in BGR, GRAY, and LAB color spaces, and comparing it with preset normal thresholds to obtain the first quality inspection result. The steps of the second quality inspection use a deep learning model (such as MobileNetV2) to predict the confidence levels of an image sample being a good image (the first confidence level) or a bad image (the second confidence level), and then obtain the second quality inspection result.
[0166] Optionally, when there is a conflict between the first quality inspection result and the second quality inspection result, that is, when the judgments of the two are inconsistent, enter the fusion decision-making stage. The above stage mainly relies on confidence level thresholds for comprehensive judgment: If the first confidence level (good image confidence level) is greater than the first confidence level threshold, it means that the deep learning model believes that the image is very likely to be a high-quality image. At this time, the second quality inspection result (the judgment of the deep learning model) is preferentially adopted, and a result of qualified quality is output. If the first confidence level is less than or equal to the first confidence level threshold, it means that the model is not confident enough in judging the image as a good image. At this time, it turns to rely on the judgment of the traditional statistical method (the first quality inspection) and outputs its result.
[0167] Optionally, a similar logic also applies to the comparison between the second confidence level (bad image confidence level) and the second confidence level threshold. Once the quality inspection result to be preferentially adopted is determined through the comparison of the confidence level and the threshold, the corresponding image quality judgment conclusion will be output, that is, qualified quality or unqualified quality. The deep learning model (such as MobileNetV2) can capture complex image features and has a high recognition ability for specific types of image defects (such as blur, color distortion, etc.); while the traditional statistical method can quickly evaluate the basic attributes of an image, such as brightness, color distribution, etc. The combination of the two can make up for the limitations of a single method and improve the accuracy and robustness of the overall detection.
[0168] Optionally, through the dynamic comparison of the confidence level and the threshold, the result of the deep learning model can be preferentially adopted when the deep learning model is confident enough, and when the deep learning model is uncertain or performs poorly, it turns to rely on the more stable traditional statistical method. The above method not only improves the accuracy of the detection but also enhances the adaptability and flexibility of the system. In the field of image quality inspection, a single evaluation criterion often fails to cover various types of image problems. The double-layer confirmation strategy proposed in this embodiment can evaluate a wider range of image quality problems by fusing the detection results of different methods, thereby improving the detection coverage and the overall quality control level.
[0169] In the embodiments of the present invention, the confidence-based image quality detection fusion decision-making strategy, through the combination of deep learning models and traditional statistical methods, and the dynamic selection according to the first confidence threshold and the second confidence threshold, can provide more accurate and robust quality detection conclusions for the original image samples. This method is particularly applicable to the pre-data preparation stage of image processing and retrieval tasks, ensuring that the images entering the subsequent processing flow meet sufficiently high quality standards, thereby improving the working efficiency and retrieval accuracy of the entire system. In a special environment such as a mine scene, the above detection method can effectively identify and eliminate low-quality images that are not suitable for image retrieval, providing high-quality data input for the image retrieval system, and thus enhancing the reliability and practicality of image queries.
[0170] As an alternative embodiment, the color information of the original image sample is determined separately in different color spaces, including: determining the mean and / or variance of the color information of the original image sample in different color spaces respectively.
[0171] In this embodiment, during the process of separately determining the color information of the original image sample in different color spaces, the mean and / or variance of the color information of the original image sample in different color spaces can be determined respectively.
[0172] Optionally, different color spaces can capture the color and brightness information of the image from different perspectives. By calculating the mean and variance of the color information in the above color spaces, the visual quality and applicability of the image can be comprehensively evaluated. This embodiment elaborates on how to extract and analyze the statistical characteristics of color information in multiple color spaces (such as BGR, GRAY, LAB) to support subsequent image quality determination.
[0173] Optionally, for the BGR channel, the original image sample is analyzed in the BGR color space. Calculate the values of the blue, green, and red color channels of each pixel point, and then calculate the mean and variance of the entire image in the BGR space. The above mean and variance can reflect the concentration degree and fluctuation range of the image color. For the GRAY space analysis, the image is converted to the GRAY space, and the mean and variance of the gray values are calculated. The analysis of the gray space mainly focuses on the brightness and contrast of the image, which helps to evaluate the clarity and richness of details of the image. The original image sample is analyzed in the LAB color space. The LAB space divides colors into brightness and two color channels, and calculates the mean and variance of the L, a, and b channels. The LAB space analysis can more meticulously evaluate the brightness, color saturation, and color balance of the image.
[0174] Optionally, after obtaining the mean and variance in different color spaces, evaluate whether these statistical characteristics fall within the preset normal range threshold. Compare the mean and variance calculated for each color space with the normal color threshold range. If the statistical values in each color space are within the normal range, it means that the color distribution and brightness level of the image are moderate, without significant color distortion or brightness abnormality. Conversely, if the statistical value of at least one color space exceeds the normal range, it indicates that the image may have problems such as color distortion, insufficient brightness or overbrightness, and abnormal contrast. The above abnormalities may affect the subsequent image processing and feature extraction processes, resulting in a decrease in the accuracy of the retrieval results.
[0175] Optionally, if the mean and variance of the color information of the original image sample in each color space are within the normal threshold range, determine that the image quality is qualified, that is, the image is suitable for subsequent feature encoding and image retrieval tasks. If the statistical value of the color information exceeds the normal threshold range in at least one color space, determine that the image quality is unqualified, that is, the image is not suitable for subsequent image processing and retrieval processes and needs to be further screened or processed.
[0176] In the embodiment of the present invention, by calculating the mean and variance of the color information in multiple color spaces respectively, a comprehensive color and brightness evaluation is provided for the original image sample. The above method can not only effectively identify image quality abnormalities such as color distortion or brightness abnormality, but also find a good balance between calculation efficiency and accuracy, and is particularly suitable for the quality control of large-scale image databases to ensure that the images used for image retrieval and feature encoding have good visual quality. In the mine scene, the above method can ensure that the images used for analysis and retrieval are of high quality and rich in information, providing a solid foundation for subsequent image processing and analysis tasks.
[0177] As an optional embodiment, perform quality detection on the original image sample to obtain a quality detection result, including: preprocess the original image sample; perform quality detection on the preprocessed original image sample to obtain a quality detection result.
[0178] In this embodiment, in the process of performing quality detection on the original image sample to obtain a quality detection result, the original image sample can be preprocessed, and then quality detection can be performed on the preprocessed original image sample to obtain a quality detection result.
[0179] Optionally, this embodiment proposes a process design for pre - processing an image and then performing quality detection. The pre - processing can improve the condition of the image, making it more suitable for the analysis of quality detection algorithms, thereby improving the accuracy and efficiency of detection. The above - mentioned process is applicable before various image - processing tasks, especially in image retrieval or feature encoding tasks, to ensure that only qualified images are used and to avoid the impact of low - quality images on the final results.
[0180] Optionally, pre - processing is an important pre - step for image quality detection. Its purpose is to eliminate or reduce noise, distortion, or irregularities in the image, making the image clearer, more regular, and easier to analyze. The pre - processing steps may include, but are not limited to: adjusting the image size to a unified size for the consistency of subsequent processing flows. Converting the image to a specific color space (such as BGR, GRAY, or LAB) and performing necessary color adjustments in that space, such as denoising, enhancing contrast, or fixing color distortion. Improving the readability of key information in the image by increasing its brightness, contrast, or sharpness. Ensuring that all images are stored in the same format for unified processing and analysis. Correcting geometric distortion caused by the lens or other factors to restore the original appearance of the image.
[0181] Optionally, pre - processing can significantly improve the accuracy of image quality detection because the pre - processed image eliminates many factors that may interfere with the detection results.
[0182] As an alternative embodiment, step S108, labeling or outputting the target image corresponding to the target feature, includes: determining the target image corresponding to the target feature; performing quality detection on the target image to obtain the quality detection result of the target image; and in response to the quality detection result of the target image indicating that the quality of the target image is qualified, labeling or outputting the target image.
[0183] In this embodiment, during the process of outputting the target image corresponding to the target feature, the target image corresponding to the target feature can be determined. Quality detection can be performed on the target image to obtain the quality detection result of the target image. If the quality detection result indicates that the quality of the target image is qualified, the target image can be output.
[0184] Optionally, this embodiment elaborates on a quality control link in the image retrieval process, mainly focusing on ensuring that the quality of the output images meets specific standards. When a user retrieves relevant images in a database through a text or image query, not only the target images related to the query need to be found, but also the quality of the above - mentioned target images must be sufficient for subsequent processing or analysis, such as feature encoding, data annotation, or visual recognition, etc. The above - mentioned method avoids the output of low - quality images by performing additional quality detection before the retrieval results are output, thereby improving the effectiveness of the entire image retrieval system and the user experience.
[0185] Optionally, after retrieving the target image, the quality of the target image is detected again to confirm whether the target image meets the predetermined quality standard. The quality detection methods used in the above process may include, but are not limited to: calculating the mean and variance in the BGR, GRAY, and LAB color spaces to ensure that the image color distribution and brightness are within the normal range. Using a trained deep learning network (such as MobileNetV2) to predict the confidence of a good image or a bad image of the image to ensure that the image is not severely distorted or damaged. A specially designed algorithm is used to identify and exclude images with screen flickering, green screen, blurring, or other visual defects.
[0186] Optionally, if the target image passes the above quality detection, that is, the quality detection result indicates that the quality of the target image is qualified, then the target image is output as a valid retrieval result to the user or integrated into the subsequent processing flow. If the quality detection result of the target image indicates that the image quality is unqualified, the target image will not be output, but continue to screen from the remaining candidate images until a qualified image is found. In some cases, if there are not enough qualified images, the user can be prompted to query again or expand the query range.
[0187] In the embodiment of the present invention, by performing additional quality detection on the target image after retrieval, dual quality control of the image retrieval result is achieved. It ensures that only truly high-quality images are used for subsequent processing, avoiding the risk of reducing the performance of the entire system due to image quality problems. The specific standards and methods of quality detection can be customized according to the specific application scenario. For example, in a mine scenario, there may be higher requirements for clarity, color authenticity, and field of view. By setting appropriate quality thresholds, it can be ensured that the output images fully meet the requirements of a specific scenario. Avoiding outputting low-quality images can significantly improve the user experience, especially in the autonomous driving system of an unmanned vehicle, where high-quality images are crucial for accurately perceiving the environment and making decisions.
[0188] As an alternative embodiment, in step S102, obtaining the input data to be queried includes: obtaining the text input data to be queried and / or the image input data to be queried.
[0189] In this embodiment, during the process of obtaining the input data to be queried, the text input data to be queried and / or the image input data to be queried can be obtained.
[0190] Optionally, this embodiment illustrates a flexible processing mechanism for user query requirements in an image retrieval system. The system is designed to accept text or images as query inputs. Thanks to deep learning models such as CLIP that can perform cross-modal feature encoding, it enables text-based image retrieval (retrieving images by text) and image-based image retrieval (retrieving images by image). The above steps ensure that the retrieval task can be flexibly executed according to different forms of query data provided by users, improving the user experience and system adaptability.
[0191] For example, users can initiate a query request by providing relevant descriptive text. For instance, in a mine scenario, users may input descriptions such as "excavator", "ore pile", or "yellow dump truck", and the system can understand the above text descriptions and convert them into queryable feature representations. The system can preprocess the input text, including word segmentation, stop word removal, conversion to lowercase, or lemmatization, etc., to improve the accuracy and efficiency of feature encoding.
[0192] For another example, users can also directly upload one or more images as query inputs, and perform retrieval based on the features of the above images. Similar to text input, images also need to be preprocessed before encoding, including size normalization, color correction, distortion removal, etc., to ensure that they are suitable for the input format and requirements of the model.
[0193] As an optional example, in some cases, users may provide both text and image inputs simultaneously. It can handle multi-modal queries, that is, simultaneously consider the feature encoding of images and text, and find the results that best match the input description and image style.
[0194] In the embodiment of the present invention, by integrating the support for text and image inputs, the flexibility and user-friendliness of the system are improved. Users can choose the most suitable query method according to the type of information at hand without worrying about the form limitations of the input data. By supporting multi-modal inputs, users can use natural language descriptions or intuitive image examples to query, which not only lowers the query threshold, improves the query accuracy, but also provides users with a more natural and intuitive interaction method.
[0195] In an embodiment of the present invention, input data of a target object in a to-be-query empty scene can be obtained, and the input data is encoded to obtain an initial feature. A target feature matching the above initial feature can be queried from the original image sample features and foreground image sample features in a database. And a target image including the target feature and the target object can be output or labeled. In this embodiment, in the face of the challenge that the target object accounts for a small proportion in the empty scene, foreground image samples having a mapping relationship with the original image samples in the empty scene can be pre-identified and extracted, and a database is established by combining the foreground image sample features and the original image sample features. In the query stage, the initial feature obtained by encoding the input data is efficiently compared with the original image sample features and foreground image sample features in the database. Even in the cold start scenario lacking specific samples, an image matching the to-be-query target object can be accurately identified, significantly improving the accuracy and efficiency of image query in the empty scene. Through the above feature encoding of the input data and database query strategy, the interference of irrelevant background information is greatly reduced, and the sensitivity to rare targets is improved, thereby ensuring the image query effect in the empty scene. The technical problem of low query accuracy of images is solved, and the technical effect of improving the query accuracy of images is achieved.
[0196] In an embodiment of the present invention, a method for training a perception model is further provided. Wherein, the method includes: obtaining input data to be queried; encoding the input data to obtain an initial feature; querying a target feature matching the initial feature from the original image sample features corresponding to the original image samples in the database and the foreground image sample features corresponding to the foreground image samples, where the original image sample is used to represent an image sample with an empty scene sample image as the background image sample and an object sample image as the foreground image sample, and the proportion of the object sample in the original image sample is less than a set threshold; labeling or outputting a target image corresponding to the target feature; training a perception model with the target image, where the perception model is used to identify the perception data collected by the driverless vehicle to obtain a perception result.
[0197] In this embodiment, input data of a target object in a to-be-query empty scene can be obtained, and the input data can be encoded to obtain initial features. Target features matching the above initial features can be queried from the original image sample features and foreground image sample features in the database. And a target image including the target features and the target object can be output or labeled. The target image obtained by the above method can be used to train a perception model, and the trained perception model can be used to identify the perception data collected by the driverless vehicle, so as to obtain a perception result. In this embodiment, in the face of the challenge that the proportion of the target object in the empty scene is small, foreground image samples having a mapping relationship with the original image samples in the empty scene can be pre-identified and extracted, and a database can be established by combining the foreground image sample features and the original image samples features. In the query stage, the initial features obtained by encoding the input data are efficiently compared with the original image sample features and the foreground image sample features in the database. Even in the cold start scenario lacking specific samples, an image matching the to-be-query target object can be accurately identified, significantly improving the accuracy and efficiency of image query in the empty scene. Through the above feature encoding of the input data and database query strategy, the interference of irrelevant background information is greatly reduced, and the sensitivity to rare targets is improved, thereby ensuring the image query effect in the empty scene. In addition, since the target image obtained by the above method can greatly reduce the interference of irrelevant background information, using the target image to train the perception model can also enable the perception model to avoid the interference of irrelevant background information and focus on identifying and analyzing the key foreground information in the target image. The perception model trained by the above method can also avoid the interference of irrelevant information in the perception data in the scenario of applying the perception data recognition, focus on the important information, and thus obtain the perception result corresponding to the important information. Furthermore, the perception effect of the driverless vehicle on the surrounding environment in the empty scene is ensured. The technical problem of low perception accuracy of driverless driving is solved, and the technical effect of improving the perception accuracy of driverless driving is achieved.
[0198] Embodiment 2
[0199] The following provides another optional specific implementation manner in combination with the image query scenario in the mine scene to elaborate on the above technical solution in detail.
[0200] Currently, the camera is a key sensor for obtaining external information. Many perception algorithms need to use the pictures or video information collected by a 2D camera, and then abstract general algorithms corresponding to the tasks therefrom. Since different algorithms are for different tasks, different types of picture data are required, which often requires screening specific data from a huge data pool. And because many tasks need cold start and lack a basic model as a discriminator, the screening task is often huge in workload and has little benefit. Therefore, there is still the technical problem of low query accuracy of images.
[0201] However, an embodiment of the present invention proposes a method for image search by image and image search by text in a mine scene based on deep learning. This method extracts the foreground image sample features that have a mapping relationship with the original image sample in an empty scene, combines the original image sample features to establish a database, realizes efficient query of the input data, and outputs the image query result. By establishing a unified feature space, the problem of inconsistent feature spaces between the image encoder and the text encoder is solved, and a vector management method is used to manage the encoding results. During query, the theoretical response speed reaches the millisecond level. It also solves the problem that abnormal encoding features caused by screen flickering and green screens in actual applications affect the query accuracy. A bad image detection module is designed, which can efficiently detect bad images with an accuracy of 96.46%. Considering the empty background and few foreground targets in the mine scene, a ROI extraction algorithm based on deep learning is designed. By encoding the features of the ROI area and establishing the association relationship between the original image and the ROI, better search capabilities for the mine scene are achieved, thereby achieving the technical effect of improving the accuracy of image query for empty scenes and solving the technical problem of low accuracy of image query.
[0202] The following is a further introduction to this method.
[0203] Figure 2 It is a flowchart of a method for image search by text and image search by image according to an embodiment of the present invention. As Figure 2 shown, the method may include the following steps:
[0204] S201, load and preprocess the image.
[0205] In this embodiment, an image to be preprocessed is obtained from the image library, and the image is loaded and preprocessed to adapt to subsequent image processing and analysis work.
[0206] Optionally, the preprocessing process may include the following steps: size normalization to ensure that the sizes of all images are consistent for easy model processing. Color conversion to convert the image into a color space suitable for model input, such as BGR or GRAY. Denoising and enhancement to remove noise in the image and enhance the clarity and contrast of the image. Distortion correction to correct the geometric distortion of the image, such as distortion or tilt. The above preprocessing process is crucial for the accuracy of image retrieval because it can eliminate irregularities in the image acquisition process and make image features more easily captured and understood by the CLIP model.
[0207] Step S202, detect the problem image.
[0208] In this embodiment, for the loaded and preprocessed images, a bad image detection algorithm (such as an algorithm based on BGR, GRAY, and LAB color space analysis) or a deep learning model, such as MobileNetV2, is used for quality assessment to identify and filter out images with a scrambled screen, a green screen, or unclear images. Excluding low-quality images can avoid interference from low-quality images on the retrieval results and ensure that the retrieved images are clear, complete, and have sufficient information content. The images after excluding low-quality images can be used as the original image samples.
[0209] Step S203, ROI extraction + ROI mapping establishment.
[0210] In this embodiment, an image detection model (such as YOLO11) is used to detect and segment the foreground objects of the preprocessed images (original image samples) to obtain the region of interest (ROI), and the images of the region of interest are used as the foreground image samples. Subsequently, a mapping relationship is established between the foreground image samples and the original image samples to ensure that subsequent feature encoding can be associated with the context information of the images. ROI extraction can focus on the key parts of the images, improving the pertinence and efficiency of feature encoding. For a mine scene with an empty background, this step is particularly important as it can help the system focus on encoding and retrieving important visual information.
[0211] Step S204, Search and configure images / texts.
[0212] In this embodiment, the query (input data) entered by the user, which can be a text description or a sample image, is received. The system prepares the query process based on the user configuration (such as whether to use ROI information, query parameters, etc.) and converts the input data into a format that can be processed subsequently. The above steps establish a bridge between the user's needs and the system's processing, ensuring that the system can perform accurate retrieval tasks based on the information provided by the user.
[0213] Step S205, Image CLIP encoding / ROI image encoding.
[0214] In this embodiment, the CLIP model is used to perform feature encoding on the processed images or ROI images, converting the image information into feature vectors for subsequent retrieval operations in the feature space. CLIP encoding is the core of cross-modal retrieval, which can ensure that images and text descriptions are in the same feature space, thus supporting the functions of image search by text and image search by image. For example, by encoding the original image samples and the foreground image samples, the corresponding original image sample features and foreground image sample features can be obtained. By encoding the input data, initial features can be obtained.
[0215] Optionally, the original image sample features and the foreground image sample features can be input into the feature library for storage.
[0216] Step S206: Calculate the similarity of the obtained features.
[0217] In this embodiment, the query feature vector (initial feature) is compared with the feature vectors of each image (or ROI image) in the feature library (original image sample feature / foreground image sample feature), and the similarity between the two is calculated, such as using cosine similarity or Euclidean distance. Similarity calculation is a key step in determining the retrieval result (search result). Based on the distance metric in the feature space, the image closest to the query feature is found.
[0218] Step S207: Determine the search result.
[0219] In this embodiment, according to the similarity calculation result, the K closest images are selected as the retrieval result and output to the user or integrated into the subsequent processing flow. Determining the search result is the end point of the retrieval process, which directly determines the quality and efficiency of the information obtained by the user, and is also the final verification of the accuracy and performance of the entire retrieval process.
[0220] In summary, during the establishment process of the feature set (feature library), image loading and preprocessing are performed on the images (image library), and the images with problems are screened, that is, problem image detection (bad image detection), and the mask of the usable part is returned. Based on the mask, the images that pass the screening are selected, and ROI extraction (ROI extraction module) is performed (to obtain ROI files), and the mapping between the ROI files and the original files (images) is established. The CLIP model is used to encode the screened original images and ROI images, and the original data feature set and the ROI image feature set are established respectively. The original data feature set and the ROI image feature set are stored in the feature library. During the processing in the query process, the input text and images are feature-encoded (CLIP encoding) to obtain encoded authentication; based on the encoded features and the above-mentioned feature library, comparison (similarity calculation) is performed, and the K nearest neighbor values are selected and returned (search result). In the overall process architecture, the ROI extraction module supports replacing different public models, and the image encoding module can support encoding using CLIP models with different numbers of parameters.
[0221] Optionally, in order to support cold start tasks, the embodiment of the present invention constructs an encoding module based on the CLIP method of zero-shot learning. The features of this method are as follows: Based on the publicly available CLIP model as a unified image and text encoding backbone (basic network architecture, responsible for extracting features from input data. For the CLIP model, its backbone part may include an image encoder and a text encoder, which are used to process image and text data respectively, and convert them into vector representations of a fixed dimension. The above vectors can be compared in a unified feature space. The encoding backbone of the CLIP model is designed to be general enough to handle diverse inputs without fine-tuning for specific tasks. This means that the same model can be applied to multiple scenarios as long as the inputs are images and texts).
[0222] Optionally, the CLIP model has the following features: Unified vector space, mapping both images and texts to the same vector space (the purpose is to make them comparable), which enables the model to directly calculate the similarity between images and texts in the vector space without additional intermediate representations. Contrastive learning, CLIP uses contrastive learning for pre-training. The model is required to map the image and text embeddings from the same sample to close positions, while mapping the (image and text) embeddings from different samples to far positions. This enables the model to learn the common features between images and texts. Multilingual support, the pre-trained model of CLIP is multilingual, which means it can process texts in multiple languages and embed them into a shared (unified) vector space. Note that this architecture uses the English version of the model because tests have found that the English version has relatively better encoding effects. Unsupervised learning, the pre-training of CLIP is unsupervised, which means it does not require a large amount of labeled data to guide the training. It learns from text and image data (in various fields) on the Internet, enabling it to perform well on tasks in various fields).
[0223] FIG. 3(a) is a schematic diagram of a problem image according to an embodiment of the present invention. As shown in FIG. 3(a), the problem image is overall blurry and lacks details, that is, a screen freeze. The feature encoding quality and matching accuracy of the blurry problem image are greatly reduced, which may lead to the retrieval results not matching the actual requirements. FIG. 3(b) is a schematic diagram of another problem image according to an embodiment of the present invention. As shown in FIG. 3(b), there is partial blurriness in this problem image, and part of it can be seen as a working vehicle in a mine scene. FIG. 3(c) is a schematic diagram of another problem image according to an embodiment of the present invention. As shown in FIG. 3(c), the entire screen is green. FIG. 3(d) is a schematic diagram of another problem image according to an embodiment of the present invention. As shown in FIG. 3(d), the upper half of the sky in the image content of this problem image is relatively clear, but the lower half of the mine scene is blurry.
[0224] Optionally, for the problem images in the data pool (image library), since the above images can show a high similarity with any input after encoding, thus affecting the query effect, therefore, before encoding the data pool, these images can be detected and excluded, that is, a screening can be done once when the data pool is established, and another such screening can also be done before the retrieval results are output.
[0225] The embodiment of the present invention sets up a bad image detection algorithm, which can achieve fast and highly accurate detection of the data pool. This algorithm determines whether there are problems with the image, such as screen freeze, green screen or blurriness, etc., by calculating the mean and standard deviation of the image in different color spaces according to a preset normal range threshold.
[0226] Optionally, this algorithm can calculate the mean and standard deviation of the BGR channels, that is, calculate the pixel mean (bgr_mean) and standard deviation (bgr_std) of the input image in the BGR color space. The BGR channels refer to the three color channels of blue (B), green (G) and red (R) in the image. By analyzing the above statistical values, it can be detected whether there are situations where the pixel values in the image are abnormally uniform or extremely dispersed, which is usually an indication of image quality problems. Convert the image to the grayscale space and calculate the mean and standard deviation. Convert the image to the grayscale (GRAY) color space and calculate the pixel mean (gray_mean) and standard deviation (gray_std) of the grayscale image. The analysis of the grayscale image helps to detect whether the overall brightness and contrast of the image are normal, which is a key indicator for judging whether the image is clear and whether there is serious noise.
[0227] Optionally, the image can also be converted to the LAB color space and channel information can be obtained. That is, the image is further converted to the LAB color space, and the mean and variance information of each channel is obtained. The design of the LAB color space is closer to human perception of colors. By analyzing the statistical values in the LAB space, it is possible to more accurately determine whether the color distribution of the image is normal, which is very effective for detecting color abnormalities (such as a flower screen or a green screen). Check whether the statistical values of each color space are within the normal range. That is, use the check_range function to check whether the pixel means and standard deviations of the BGR, GRAY, and LAB color spaces fall within the preset normal range. Judging whether the statistical characteristics of the image in different color spaces meet the standards of high-quality images helps to exclude those images that exhibit abnormal characteristics in any color space. Comprehensively judge whether there are problems with the image. If the statistical values of the image in the three color spaces of BGR, GRAY, and LAB all meet the normal range, the algorithm returns False, indicating that the image quality is normal; on the contrary, if the statistical value of any space exceeds the normal range, the algorithm returns True, indicating that there is a problem with the image. Ensure that the image is considered of high quality only when the analysis results in each color space are normal, which improves the accuracy and robustness of the detection of problematic images.
[0228] In the embodiment of the present invention, by analyzing the statistical characteristics of the image in the three color spaces of BGR, GRAY, and LAB, the algorithm can comprehensively detect various quality problems that the image may have, including color abnormalities, uneven brightness, and noise. Using the mean and standard deviation as indicators for image quality evaluation, these statistical characteristics can concisely and effectively describe the pixel value distribution of the image, facilitating the setting of thresholds for automated judgment. By setting the normal range thresholds (bgr_range, gray_range, lab_range) in different color spaces, the algorithm can flexibly adapt to the quality requirements of different scenarios and image types, ensuring the accuracy of the detection of problematic images. The computational cost of calculating the mean and standard deviation is relatively low, enabling the algorithm to quickly process a large number of images, suitable for the rapid screening of image datasets, and improving the efficiency of the image retrieval system.
[0229] Optionally, by calculating the color means and variances of the input image in different spaces and setting reasonable thresholds based on the statistical values, a detection accuracy of 96.46% for detecting bad images is achieved. In addition, the embodiment of the present invention also uses the mobilenetV2 deep learning method for identifying bad images, outputting the bad image confidence and prediction probability of each image; when the output of mobilenetV2 is inconsistent with the output result of the traditional algorithm proposed in the present invention, cross-validation is performed. If the confidence of the mobilenetV2 network is higher than the threshold during cross-validation then it is considered that there is no problem with the image. If the confidence is lower than the threshold Then it is considered that the image is a bad image. (When the two algorithms are inconsistent, the result of mobilenetV2 will be further judged: when the confidence of mobilenetV2 that the above image is a good image or a bad image is very high (greater than 0.85), it is considered a good image or a bad image; when the confidence of mobilenetV2 does not meet the aforementioned threshold confidence, the result will be output according to the traditional method.)
[0230] By calculating the color mean and variance of the input image in different spaces and setting reasonable thresholds according to the statistical values, the detection accuracy of 96.46% for bad image detection is achieved; by fusing the results of mobilenetV2, the recognition accuracy is increased to 98.5%. Based on the evaluation and analysis of the embodiments of the present invention, it is found that the bad images that cannot be detected generally have small damaged areas, only part of them are damaged, and the impact on the query results during query is also small. Based on the fact that the bad images that cannot be detected generally have small damaged areas, only part of them are damaged, and the impact on the query results during query is also small.
[0231] Figure 4 It is a schematic diagram of a network structure that can be viewed only once according to an embodiment of the present invention, as Figure 4 shown. The YOLO11 network structure includes a backbone network 401, a neck network 402 (Neck), and a head network 403 (Head). Among them, the backbone network 401 may include a Channel-to-Point Spatial Attention module (abbreviated as C2PSA), a Spatial Pyramid Pooling with Fixed size bins (abbreviated as SPPF), a Convolution module with 3 layers using a kernel size of 2x2 (abbreviated as C3K2), a Convolution + Batch Normalization + Swish activation function (abbreviated as CBS). The neck network 402 may include an Upsample, a feature map Contact, a Convolution module with 3 layers using a kernel size of 2x2, a Convolution + Batch Normalization + Swish activation function, and other modules. The head network 403 may include CBS, a Depthwise Separable Convolution (abbreviated as DSC), and a 2D Convolution (abbreviated as Con2d).
[0232] Optionally, in order to improve query efficiency, the embodiment of the present invention establishes a feature library based on the Faiss database. Faiss is an efficient similarity retrieval and clustering engine for dense vectors, which can create a millisecond-level nearest neighbor search on a billion-level data set.
[0233] In this embodiment, a method for searching images by text and searching images by images is designed for mining scenes. A unified feature map of images and texts is mainly established through the CLIP model (not limited to the CLIP model). The encoding and mapping accuracy of this method is very low when the background is empty. The foreground is extracted first and then the mapping is performed to improve the accuracy. A feature library is established based on faiss, and then the results of searching images by text and searching images by images are returned by performing a similarity query by performing CLIP feature encoding on the input text and image. Targeted algorithm detection is designed for the problems of flower screen and green screen encountered in the actual process, and bad images are detected with a detection accuracy of 96.46%. In addition, considering that the background of the mining scene is empty, an ROI extraction module is designed, and the foreground in the scene is selected as the ROI area and coded with emphasis, finally realizing an efficient method of searching images by text and searching images by images.
[0234] FIG5(a) is a schematic diagram of search results obtained by searching in a text-to-image manner according to an embodiment of the present invention. As shown in FIG5(a), the search text is human and the region of interest is not used, that is, the user's search text is "searchtext: input: human" and the ROI is not used to obtain the three images containing human beings shown. Among them, the image on the left side of FIG5(a) contains human being 501, the image in the middle of FIG5(a) contains human being 502, and the image on the right side of FIG5(a) contains human being 503.
[0235] FIG5(b) is a schematic diagram of another search result obtained by searching in a text-to-image manner according to an embodiment of the present invention. As shown in FIG5(b), the search text is a person and the region of interest is not used, that is, the user's search text is "searchtext: input: human", and the ROI can be used to obtain the three types of images containing people shown. Among them, the image on the left side of FIG5(b) contains a person 51, the image in the middle of FIG5(b) contains a person 52, and the image on the right side of FIG5(b) contains a person 53.
[0236] FIG. 5(c) is a schematic diagram of another search result obtained by the text-based image search method according to an embodiment of the present invention. As shown in FIG. 5(c), the retrieved text is water and the region of interest is not used. That is, the user's search text is "searchtext: input: water", and the three shown pictures containing ponding water can be obtained without using the ROI. Among them, the left picture in FIG. 5(c) contains ponding water 54, the middle picture in FIG. 5(c) contains ponding water 55, and the right picture in FIG. 5(c) contains ponding water 56.
[0237] FIG. 5(d) is a schematic diagram of another search result obtained by the text-based image search method according to an embodiment of the present invention. As shown in FIG. 5(d), the retrieved text is water and the region of interest is used. That is, the user's search text is "searchtext: input: water", and the three shown pictures containing ponding water can be obtained by using the ROI. Among them, the left picture in FIG. 5(d) contains ponding water 57, the middle picture in FIG. 5(d) contains ponding water 58, and the right picture in FIG. 5(d) contains ponding water 59.
[0238] FIG. 5(e) is a schematic diagram of another search result obtained by the text-based image search method according to an embodiment of the present invention. As shown in FIG. 5(e), the retrieved text is human and white truck and the region of interest is not used. That is, the user's search text is "search text: input: human with white truck", and the three shown pictures containing both human and white truck can be obtained without using the ROI. Among them, the left picture in FIG. 5(e) contains human and white truck 60, the middle picture in FIG. 5(e) contains human and white truck 61, and the right picture in FIG. 5(e) contains human and white truck 62.
[0239] FIG. 5(f) is a schematic diagram of another search result obtained by the text-based image search method according to an embodiment of the present invention. As shown in FIG. 5(f), the retrieved text is human and white truck and the region of interest is used. That is, the user's search text is "search text: input: human with white truck", and the three shown pictures containing ponding water can be obtained by using the ROI. Among them, the left picture in FIG. 5(f) contains human and white truck 63, the middle picture in FIG. 5(f) contains human and white truck 64, and the right picture in FIG. 5(f) contains human and white truck 65.
[0240] FIG. 6(a) is a schematic diagram of an input picture of the picture to be searched by the image-based image search method according to an embodiment of the present invention. As shown in FIG. 6(a), it can be an input picture containing staff and vehicles to be subjected to image search in the user input system. Among them, the picture in FIG. 6(a) contains human and vehicle 601.
[0241] FIG. 6(b) is a schematic diagram of search results obtained by searching based on an input image in a manner of searching images by image according to an embodiment of the present invention. As shown in FIG. 6(b), taking the above FIG. 6(a) as the input image and without using the region of interest, that is, the user inputs "search images: input: None" and does not use the ROI, the three search results shown can be obtained. Among them, the left image in FIG. 6(b) contains a person and a vehicle 66, the middle image in FIG. 6(b) contains a person and a vehicle 67, and the right image in FIG. 6(b) contains a person and a vehicle 68.
[0242] FIG. 6(c) is a schematic diagram of another search result obtained by searching based on an input image in a manner of searching images by image according to an embodiment of the present invention. As shown in FIG. 6(c), taking the above FIG. 6(a) as the input image and using or not using the region of interest, that is, the user inputs "search images: input: None" and uses the ROI, the three search results shown can be obtained. Among them, the left image in FIG. 6(c) contains a person and a vehicle 69, the middle image in FIG. 6(c) contains a person and a vehicle 70, and the right image in FIG. 6(c) contains a person and a vehicle 71.
[0243] As can be seen from the above, after using the ROI information in the embodiment of the present invention, the system can better process the focus and details of the image, improving the efficiency, accuracy and robustness of the search task, especially in an environment such as a mine that requires high-precision identification of specific objects or elements. By specifically encoding and processing the ROI, the performance of cross-modal search can be significantly improved.
[0244] In an embodiment of the present invention, input data of a target object in a to-be-query empty scene can be obtained, and the input data can be encoded to obtain an initial feature. A target feature matching the above initial feature can be queried from the original image sample features and foreground image sample features in a database. And a target image including the target feature and the target object can be output or labeled. In this embodiment, in the face of the challenge that the target object accounts for a small proportion in an empty scene, foreground image samples having a mapping relationship with the original image samples in the empty scene can be pre-identified and extracted, and a database can be established by combining the foreground image sample features and the original image sample features. In the query stage, the initial feature obtained by encoding the input data is efficiently compared with the original image sample features and the foreground image sample features in the database. Even in a cold start scenario lacking specific samples, an image matching the to-be-query target object can be accurately identified, significantly improving the accuracy and efficiency of image query in an empty scene. Through the above feature encoding of input data and database query strategy, the interference of irrelevant background information is greatly reduced, and the sensitivity to rare targets is improved, thereby ensuring the image query effect in an empty scene. The technical problem of low query accuracy of images is solved, and the technical effect of improving the query accuracy of images is achieved.
[0245] Embodiment 3
[0246] An embodiment of the present invention provides an image query device. It should be noted that the image query device in the embodiment of the present invention can be used to execute Figure 1 the image query method provided in the embodiment of the present invention herein. The following introduces the image query device provided in the embodiment of the present invention.
[0247] Figure 7 is a schematic structural diagram of an image query device according to an embodiment of the present invention. As Figure 7 shown, the device may include: an acquisition unit 702, an encoding unit 704, a query unit 706, and a processing unit 708.
[0248] The acquisition unit 702 is configured to acquire input data to be queried.
[0249] The encoding unit 704 is configured to encode the input data to obtain an initial feature.
[0250] The query unit 706 is configured to query a target feature matching the initial feature from the original image sample features corresponding to the original image samples in the database and the foreground image sample features corresponding to the foreground image samples.
[0251] The processing unit 708 is configured to label or output a target image corresponding to the target feature.
[0252] The image query device provided by the embodiment of the present invention obtains the input data to be queried through the obtaining unit 702; encodes the input data through the encoding unit 704 to obtain the initial features; queries for the target features matching the initial features in the original image sample features corresponding to the original image samples in the database and the foreground image sample features corresponding to the foreground image samples through the query unit 706; and labels or outputs the target images corresponding to the target features through the processing unit 708, thereby solving the technical problem of low query accuracy of images and achieving the technical effect of improving the query accuracy of images.
[0253] The above device may further include a processor and a memory. The above units are all stored in the memory as program units, and the corresponding functions are implemented by the processor executing the above program units stored in the memory.
[0254] The above processor includes a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and the graceful shutdown of the devices to be shut down of the same device type is controlled by adjusting the kernel parameters.
[0255] The above memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (Random Access Memory, abbreviated as RAM) and / or non-volatile memory, such as read-only memory (Read-Only Memory, abbreviated as ROM) or flash RAM (flash RAM), and the memory includes at least one memory chip.
[0256] The processor includes a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and the work efficiency of traders is improved by adjusting the kernel parameters.
[0257] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM (flash RAM), and the memory includes at least one memory chip.
[0258] Embodiment 4
[0259] According to the embodiment of the present invention, there is also provided a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the above method is implemented.
[0260] Embodiment 5
[0261] According to the embodiment of the present invention, there is also provided a processor, and the processor is used to run a program, wherein when the program runs, the above method is executed.
[0262] Embodiment 6
[0263] Figure 8 is a schematic diagram of an electronic device for an image query method according to an embodiment of the present invention. As Figure 8 shown, an embodiment of the present invention further provides an electronic device 800. The device includes a processor 801, a memory 802, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the method in any of the above embodiments, which will not be elaborated here.
[0264] The devices herein can be servers, personal computers (PCs for short), personal access devices (PADs for short), mobile phones, etc.
[0265] Embodiment 7
[0266] The present invention also provides a computer program product, which is adapted to execute a program initialized with any of the above method steps when executed on a data processing device.
[0267] Embodiment 8
[0268] According to another aspect of an embodiment of the present invention, a vehicle is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute any of the method steps.
[0269] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
[0270] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0271] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for querying an image, characterized in that, Including: Obtain input data to be queried, where the input data is used to represent a target object in an empty scene to be queried; Encode the input data to obtain an initial feature; In the original image sample features corresponding to the original image samples in the database and the foreground image sample features corresponding to the foreground image samples, query for target features that match the initial feature, where there is a mapping relationship between the original image sample and the foreground image sample, the original image sample is used to represent an image sample with an empty scene sample image as the background image sample and an object sample image as the foreground image sample, and the proportion of the size of the object sample in the original image sample is less than a set threshold; Label or output the target image corresponding to the target feature, where the target image includes the target object.
2. The method according to claim 1, characterized in that, Querying for target features that match the initial feature in the original image sample features corresponding to the original image samples in the database and the foreground image sample features corresponding to the foreground image samples includes: In the original image sample features in the database, query for original image sample features whose similarity to the initial feature is greater than a similarity threshold, and / or, in the foreground image sample features in the database, query for foreground image sample features whose similarity to the initial feature is greater than the similarity threshold; Determine the queried original image sample features and / or the queried foreground image sample features as the target features that match the initial feature.
3. The method according to claim 1, characterized in that, The method further includes: Perform foreground extraction on the original image sample to obtain the foreground image sample; Establish the mapping relationship between the foreground image sample and the original image sample; According to the mapping relationship, store the foreground image sample features and the original image sample features in the database.
4. The method according to claim 3, characterized in that, Performing foreground extraction on the original image sample to obtain the foreground image sample includes: Perform foreground detection on the original image sample to obtain detection information; Based on the detection information, extract the foreground image sample from the original image sample.
5. The method according to claim 4, characterized in that, Performing foreground detection on the original image sample to obtain detection information includes: Call an information extraction model to perform foreground detection on the original image sample to obtain the detection information, where the information extraction model is trained using detection information samples of object samples in the empty scene; Based on the detection information, extracting the foreground image sample from the original image sample includes: using the information extraction model to analyze the detection information and the original image sample to obtain the foreground image sample.
6. The method according to claim 3, characterized in that, Storing the foreground image sample features and the original image sample features in the database according to the mapping relationship includes: Perform feature encoding on the foreground image sample and the original image sample with the mapping relationship in the target feature space respectively to obtain the foreground image sample features and the original image sample features with the mapping relationship; Store the foreground image sample features and the original image sample features with the mapping relationship in the database; Encode the input data to obtain initial features, including: performing feature encoding on the input data in the target feature space to obtain the initial features.
7. The method according to claim 3, characterized in that, The method further includes: Perform quality detection on the original image sample to obtain the quality detection result of the original image sample; Extract the foreground from the original image sample to obtain the foreground image sample, including: in response to the quality detection result indicating that the quality of the original image sample is qualified, extract the foreground from the original image sample to obtain the foreground image sample.
8. The method according to claim 7, characterized in that, The quality detection result includes a first quality detection result. Performing quality detection on the original image sample to obtain the quality detection result of the original image sample includes: Respectively determine the color information of the original image sample in different color spaces to obtain multiple color information; In response to each of the multiple color information being within the normal color threshold range, determine that the first quality detection result is that the quality of the original image sample is qualified; In response to at least one of the multiple color information not being within the normal color threshold range, determine that the first quality detection result is that the quality of the original image sample is unqualified.
9. The method according to claim 8, characterized in that, The quality detection result includes a second quality detection result. Performing quality detection on the original image sample to obtain the quality detection result of the original image sample includes: Determine the first confidence level of the original image sample, where the first confidence level is used to represent the degree of possibility that the quality of the original image sample is qualified; in response to the first confidence level being greater than the first confidence level threshold, determine that the second quality detection result is that the quality of the original image sample is qualified; in response to the first confidence level being less than or equal to the first confidence level threshold, determine that the second quality detection result is that the quality of the original image sample is unqualified; Or, Determine the second confidence level of the original image sample, where the second confidence level is used to represent the degree of possibility that the quality of the original image sample is unqualified; in response to the second confidence level being greater than the second confidence level threshold, determine that the second quality detection result is that the quality of the original image sample is unqualified; in response to the second confidence level being less than or equal to the second confidence level threshold, determine that the second quality detection result is that the quality of the original image sample is qualified.
10. The method according to claim 9, characterized in that, In the case where the first quality detection result is different from the second quality detection result, the method further includes: In response to the first confidence level being greater than the first confidence level threshold, output the second quality detection result; in response to the first confidence level being less than or equal to the first confidence level threshold, output the first quality detection result; or In response to the second confidence level being greater than the second confidence level threshold, output the second quality detection result; in response to the second confidence level being less than or equal to the second confidence level threshold, output the first quality detection result.
11. The method according to claim 8, characterized in that, Determine the color information of the original image sample in different color spaces respectively, including: Determine the mean value and / or variance of the color information of the original image sample in different color spaces respectively.
12. The method according to any one of claims 1 to 11, characterized in that, Label or output the target image corresponding to the target feature, including: Determine the target image corresponding to the target feature; Perform quality detection on the target image to obtain the quality detection result of the target image; In response to the quality detection result of the target image indicating that the quality of the target image is qualified, label or output the target image.
13. A vehicle, characterized in that, Including: A memory storing an executable program; A processor for running the program, wherein when the program runs, it executes the method according to any one of claims 1 to 12.
14. An electronic device, characterized in that, Including: A memory storing an executable program; A processor for running the program, wherein when the program runs, it executes the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Image color adjustment method and device based on scene, and computer equipment
CN111062860A
Image feature storage method, image query method and related devices
CN118152608A
Target detection method and device, terminal equipment and storage medium
CN118262208A
Wireless Mouse
KR1020230072147A