Image query method, device, and program product
By extracting traditional and deep features in image materials and generating fusion features for image matching, the problem of time-consuming and error-prone image query in the prior art is solved, and efficient and accurate automated image query is achieved.
Patent Information
- Application Number
- CN202410995937.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-07-24
AI Technical Summary
The existing image query methods are time-consuming and labor-intensive, have low query efficiency and are prone to errors, and have poor accuracy.
By obtaining the image material corresponding to the image to be queried, traditional features and depth features are extracted, fusion features are generated, and image matching is performed based on the fusion features to display the matching target image.
It realizes automated image query, improves query efficiency and accuracy, and reduces omissions and errors.
Smart Images

Figure CN119088997B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image search technology. Specifically, this application relates to an image query method, device, and program product. Background Art
[0002] With the development of the electronic technology field and the progress of image shooting technology, people have generated a large amount of image materials through various image shooting devices (such as mobile phones, cameras, cameras, etc.). These image materials often contain a large amount of valuable image information. For example, when a user loses an item in a public place equipped with monitoring devices, the monitoring video output by the monitoring device can be analyzed and queried based on the relevant information about the item provided by the user, and the image information of the item can be found in the monitoring video to facilitate the quick search for the item.
[0003] Currently, the query of image information in image materials still relies on the method of manually retrieving each image one by one. This query method not only takes time and effort, has low query efficiency, but also is prone to errors, reducing the accuracy of image query. Summary of the Invention
[0004] Embodiments of this application provide an image query method, device, and program product, which can solve the problems of the existing image query method being time-consuming and laborious, having low query efficiency, being prone to errors, and having poor accuracy. To achieve this purpose, the embodiments of this application provide the following several solutions.
[0005] According to one aspect of the embodiments of this application, an image query method is provided, including:
[0006] Obtain the image material corresponding to the image to be queried, where the image material includes at least one of video and picture;
[0007] Extract the traditional features and depth features in the image to be queried and the image material according to a preset scale, and generate the fusion features corresponding to the traditional features and depth features;
[0008] Based on the fusion features and the preset scale, perform feature matching between the image to be queried and the image material;
[0009] According to the matching result, display the target image in the image material that matches the image to be queried.
[0010] In a possible implementation manner, the obtaining the image material corresponding to the image to be queried includes:
[0011] Obtain the image to be queried according to the obtained image capture operation;
[0012] If it is determined that the current situation meets the preset conditions, obtain the image material corresponding to the image to be queried according to the preset conditions, where the preset conditions include at least one of image material import, receiving a search instruction, and currently loading image material.
[0013] In a possible implementation manner, obtaining the image to be queried according to the obtained image capture operation includes:
[0014] Load the video file and play the video frame corresponding to the video file;
[0015] Determine the cropping area in the video frame according to the image capture operation of the cropping object, and crop the video frame based on the cropping area to obtain the image to be queried.
[0016] In a possible implementation manner, extracting the traditional features and depth features from the image to be queried and the image material according to the preset scale includes:
[0017] Construct an image pyramid corresponding to the image to be queried and the image material based on the preset scale, where the preset scale includes at least one of object size and image resolution;
[0018] Extract the traditional features and the depth features at each scale of the image pyramid.
[0019] In a possible implementation manner, generating the fusion features corresponding to the traditional features and the depth features includes:
[0020] Obtain the current weights corresponding to the traditional features and the depth features, and fuse the traditional features and the depth features based on the current weights to generate the fusion features, where the current weights are generated based on the initial weights, the current scene, the current scale, and the error feedback.
[0021] In a possible implementation manner, performing feature matching between the image to be queried and the image material based on the fusion features and the preset scale includes:
[0022] Match the fusion features of the image to be queried with the fusion features of the image material at each scale corresponding to the preset scale;
[0023] Obtain the matching result between the image to be queried and the image material according to the matching result corresponding to each scale.
[0024] In a possible implementation manner, obtaining the matching result between the image to be queried and the image material according to the matching result corresponding to each scale includes:
[0025] Obtain matching data within a preset time period according to the matching result, where the matching data includes the quantity and quality of the matching result;
[0026] Adjust the extraction parameters corresponding to the traditional features and the deep features based on the matching data, where the extraction parameters include the feature extraction density.
[0027] In a possible implementation manner, the image material is a video. After obtaining the matching result between the image to be queried and the image material according to the matching result corresponding to each scale, it includes:
[0028] Obtain the matching feature points in each frame according to the matching result;
[0029] Track the motion trajectory of the feature points according to the timing information of the video, and filter the feature points based on the motion trajectory.
[0030] According to one aspect of the embodiments of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the method as described above.
[0031] According to one aspect of the embodiments of the present application, there is provided a computer program product, including a computer program, where when the computer program is executed by a processor, the steps of the method as described above are implemented.
[0032] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are:
[0033] The image query method provided by the present application obtains the image material corresponding to the image to be queried; extracts the traditional features and deep features in the image to be queried and the image material according to the preset scale, and generates the fusion features corresponding to the traditional features and the deep features; performs feature matching between the image to be queried and the image material based on the fusion features and the preset scale; and displays the target image in the image material that matches the image to be queried according to the matching result. The embodiments of the present application extract the traditional features and deep features in the image to be queried and the image material, generate fusion features by combining the traditional features and the deep features, perform feature matching between the query image and the image material according to the fusion features and the preset scale, and obtain the target image in the image material corresponding to the image to be queried. The embodiments of the present application can automatically search for the target image in the image material that matches the image to be queried, the query method is simple and efficient, and it is not easy to miss and make mistakes, effectively improving the accuracy of image query. Description of the Drawings
[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application.
[0035] Figure 1Flowchart of the image query method provided by the embodiments of the present application;
[0036] Figure 2 Flowchart of image matching in the image query method provided by the embodiments of the present application;
[0037] Figure 3 Workflow diagram of the image query method provided by the embodiments of the present application;
[0038] Figure 4 Flowchart of image capture provided by the embodiments of the present application;
[0039] Figure 5 Structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0040] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0041] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude being implemented as other features, information, data, steps, operations, elements, components, and / or combinations thereof supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein can include wireless connection or wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" indicates being implemented as "A", or being implemented as "B", or being implemented as "A and B".
[0042] To make the purpose, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings.
[0043] The technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0044] The image query method, device, and program product provided by this application aim to solve at least one technical problem existing in the prior art.
[0045] In an embodiment of this application, an image query method is provided. The device that executes this image query method can be a mobile phone, a server, a laptop computer, a monitoring device management platform, and other devices that can obtain the image to be queried and the image materials to be matched.
[0046] Optionally, as Figures 1 - 4 shown, this image query method includes:
[0047] S101: Obtain the image materials corresponding to the image to be queried.
[0048] Optionally, the image materials include at least one of video and picture. Obtaining the image materials corresponding to the image to be queried includes: obtaining the image to be queried according to the obtained image capture operation; if it is determined that the current meets the preset conditions, obtaining the image materials corresponding to the image to be queried according to the preset conditions, and the preset conditions include at least one of image material import, receiving a search instruction, and currently loading image materials.
[0049] Optionally, it is also possible not to obtain the image to be queried through the image capture operation, and the imported image to be queried can be obtained according to the user's import operation or the image to be queried can be determined according to the instruction indicating the image to be queried input by the user.
[0050] Optionally, after obtaining the image to be queried, the image to be queried can be saved, and for the convenience of the user to view the image to be queried, after obtaining the image to be queried, a new window can also be displayed, and the image to be queried is displayed in the new window. Among them, for the convenience of the user to view and discover the image to be queried, the image to be queried can also be highlighted.
[0051] Optionally, when the preset condition is image material import, the imported image materials can be determined as the image materials corresponding to the image to be queried; when the preset condition is receiving a search instruction, the currently loaded image materials can be determined as the image materials corresponding to the image to be queried, or the image materials corresponding to the query instruction can be obtained according to the information in the search instruction or the image material selection instruction related to the search instruction; when the preset condition is that image materials are currently loaded, the currently loaded image materials can be directly determined as the image materials corresponding to the image to be queried.
[0052] Optionally, obtaining the image to be queried according to the acquired image capture operation includes: loading a video file and playing the video frame corresponding to the video file; determining a cropping area in the video frame according to the image capture operation of the capture object, and cropping the video frame based on the cropping area to obtain the image to be queried. Among them, the capture object can be a user, and the cropping area in the video frame is determined by the image capture operation performed by the user.
[0053] Optionally, the image capture operation may include an image capture tool selection operation and an image capture tool movement operation. According to the image capture tool selection operation, determine the image capture tool for capturing the image, and move the image capture tool to the image screen based on the image capture tool movement operation, and determine the area occupied by the image capture tool in the video frame as the cropping area. It is also possible to obtain human-computer interaction information (such as the movement trajectory of the mouse on the video screen, the finger sliding trajectory, etc.) through the image capture operation, and determine the cropping area according to the human-computer interaction information.
[0054] In one embodiment, use the multimedia module based on Qt or a third-party library (such as FFmpeg) to read the video file, and convert the video frames in the video file into the image format of Qt. Use the image display control of Qt (such as QLabel) to display the video frames on the interface so that the user can perform operations to crop the screen. The user can perform an image capture operation through other methods such as mouse interaction, draw a rectangle, polygon or custom area of any shape on the video screen, and use the drawn area as the cropping area. Perform corresponding image processing operations (such as cropping, masking, etc.) on the displayed video frames based on the cropping area to obtain the cropped screen required by the user. Display the cropped screen on the interface or save it as a file for subsequent processing. A graphic toolbar can also be displayed on this interface. The graphic toolbar includes graphic tools of multiple shapes, and the cropping area is circled in the video screen through the graphic tools.
[0055] S102: Extract the traditional features and depth features in the image to be queried and the image material according to the preset scale, and generate the fusion features corresponding to the traditional features and depth features.
[0056] Optionally, when the image material is a video, before extracting the traditional features and depth features, first extract video frames from the video, and perform preprocessing on each video frame. The preprocessing includes at least one of grayscale conversion, denoising, and smoothing processing.
[0057] Optionally, extracting the traditional features and depth features in the image to be queried and the image material according to the preset scale includes: constructing an image pyramid corresponding to the image to be queried and the image material based on the preset scale, and the preset scale includes at least one of object size and image resolution; extracting traditional features and depth features at each scale of the image pyramid.
[0058] Optionally, the traditional features include feature points extracted by methods such as SIFT and SURF. The descriptor of the traditional feature is a vector describing the local area of the feature point. The depth features can be extracted by a pre-trained convolutional neural network (CNN), and the extracted feature points are represented in the form of feature vectors.
[0059] In one embodiment, the image material is a vehicle image. Feature points such as headlight, license plate, and window corner points and their descriptors are extracted using SIFT. Among them, for the feature points at the position of the front headlight, the descriptor contains the gradient information of the area around the front headlight. For the feature points at the edge of the wheel, the descriptor contains the information of the texture change of the wheel. For the feature points at the window corner, the descriptor contains the gradient direction and amplitude of the window edge. For the feature points at the position of the door handle, the descriptor contains the image features around the door handle, and high-level semantic features of the body contour, headlight shape, and license plate area extracted by a pre-trained CNN.
[0060] Optionally, fusion features corresponding to the traditional features and the depth features are generated, including: obtaining the current weights corresponding to the traditional features and the depth features, and fusing the traditional features and the depth features based on the current weights to generate fusion features. The current weights are generated based on the initial weights, the current scene, the current scale, and the error feedback.
[0061] Optionally, the initial weights of the traditional features and the depth features can both be 0.5. The importance and similarity of the depth features and the traditional features can be determined according to the current scene, the current scale, and the error feedback. Based on the importance and similarity, the current weights are determined, and then the traditional features and the depth features are weighted and fused according to the current weights.
[0062] Optionally, algorithms such as Adaptive Boosting and Meta-Learning can be used to dynamically adjust the weights in real time. The error feedback information includes the accuracy and stability of the matching results. The error feedback information can be obtained based on the matching results of other objects in the same image material or the matching results of other image materials in the same scene. The above algorithms are trained based on the historical data of the scene, scale, and error feedback, and the trained algorithms are used for feature fusion. When performing feature fusion, the algorithms for adjusting the weights determine the current weights of the traditional features and the depth features according to the error feedback, the current scene, and the current scale, and perform feature fusion based on the current weights.
[0063] Optionally, the performance of traditional features and depth features during image matching can be obtained, and error feedback can be generated based on the performance results. Specifically, if traditional features (such as SIFT) show higher accuracy during image query within several consecutive frames (it is easier to query and match the target image through traditional features), then error feedback is generated to increase the weight of traditional features and decrease the weight of depth features. Conversely, if the performance of depth features (such as the body contour recognized by CNN) is better, then error feedback is generated to increase the weight of depth features, thereby improving the adaptability of the algorithm.
[0064] In one embodiment, the image material is a video. Different-scale image pyramids are constructed based on the query image and the video frames. Each pyramid level represents a different scale, such as the original image, an image reduced to one-half of the original image, an image reduced to one-third of the original image, and so on. During the feature extraction process, traditional features and depth features are respectively extracted for each pyramid level. For example, for the original image, feature points and descriptors of traditional features are extracted; for the image reduced to one-half of the original image, feature points and descriptors of traditional features are extracted; for the image reduced to one-third of the original image, feature points and descriptors of traditional features are extracted. By the above method, features can be detected and described at different scales. And the traditional features and depth features are fused at each scale to obtain fused features.
[0065] S103: Perform feature matching between the query image and the image material based on the fused features and a preset scale.
[0066] Optionally, performing feature matching between the query image and the image material based on the fused features and a preset scale includes: matching the fused features of the query image with the fused features of the image material at each scale corresponding to the preset scale; obtaining the matching result between the query image and the image material according to the matching results corresponding to each scale, where the fused features can be vectors.
[0067] Optionally, feature matching can be performed by means of distance measurement, and the distance measurement includes Euclidean distance, cosine similarity, etc. Determine the correspondence between the fused features in the query image and the fused features in each image material (such as each video frame), generate feature pairs according to the correspondence, and calculate the distance between the two fused features in the feature pairs. Based on the average distance or weighted distance of all feature pairs at the current scale, the matching score at the current scale is obtained. The multi-scale matching method of this application can process images or video streams with large-scale changes, improve the accuracy and stability of matching, and can adapt to scenarios where image resolutions are inconsistent or the sizes of target objects change.
[0068] Optionally, corresponding weights can be preset for the matching results of different scales. The matching result between the image to be queried and the image material can include a matching score, which is obtained by weighting based on the weight. And the weights corresponding to the matching results of different scales can be adjusted according to matching feedback (such as at least one of the stability and accuracy of the matching score, and the received parameter adjustment instruction). When the image material is a video, the matching result can further include the timestamp and position information of the matching frame (such as the video frame with a matching score greater than a preset threshold) determined based on the matching score, and the obtained matching result is saved for subsequent analysis and processing.
[0069] Optionally, when the image material is a video, to eliminate accidental matching, after obtaining the matching result between the image to be queried and the image material according to the matching results corresponding to each scale, it includes: obtaining the matching feature points in each frame according to the matching result; wherein, the traditional features and depth features in the video frame are represented in the form of feature points; tracking the motion trajectory of the feature points according to the timing information of the video, and screening the feature points based on the motion trajectory. Performing temporal consistency analysis on the feature points using the timing information can better identify the targets in the dynamic scene and obtain the motion trajectory of the targets, and eliminate the feature points that affect the matching result.
[0070] In one embodiment, the timing information includes the timestamp or sequence number of each video frame, which is used to determine the chronological order and interval of each frame in time. Obtain the feature points in each video frame that match the image to be queried, and calculate and determine the motion trajectory of the feature points in the video according to the timing information. Calculate the displacement or speed of the feature points according to the motion trajectory, and screen the inconsistent or abnormal (such as the displacement distance between adjacent frames is greater than a preset distance or the speed is greater than a preset speed) feature points based on the displacement or speed, so as to improve the accuracy and stability of the matching result. Among them, the operation of eliminating inconsistent or abnormal feature points through the timing information can also be performed after feature point extraction. After extracting the feature points of each video frame in the video, extract the inconsistent or abnormal feature points in the video frame according to the timing information.
[0071] Optionally, obtaining the matching result between the image to be queried and the image material according to the matching results corresponding to each scale includes: obtaining the matching data within a preset time period according to the matching result, and the matching data includes the quantity and quality of the matching results; adjusting the extraction parameters corresponding to the traditional features and depth features based on the matching data, and the extraction parameters include the feature extraction density.
[0072] Optionally, the quantity of the matching result can be the number of matching feature points in the image material, and the quality can be the comparison result between the number of matching feature points and a predetermined threshold.
[0073] In one embodiment, the extraction parameter is sift->getEdgeThreshold(), and the feature point detection sensitivity is changed by adjusting this extraction parameter. The matching data includes the matchCount parameter, which is used to indicate the number of matching results in the current frame or within a certain time period. Based on this number, the feature point matching density in the current scene can be inferred, and thus the extraction parameter can be dynamically adjusted. When the matchCount parameter indicates a small number (e.g., less than a predetermined threshold), it means that the number of feature point matches is insufficient, and the extraction amount of matching points can be increased by increasing the feature extraction density or adjusting the relevant threshold. When the matchCount parameter indicates a large number, it means that the number of feature point matches is excessive, which may introduce noise or unnecessary feature points. The number of matching feature points can be dynamically and adaptively adjusted by reducing the feature extraction density or adjusting the relevant threshold.
[0074] Optionally, each piece of material (such as a photo or video frame) in the image material is matched with the object to be queried until the object to be queried is completely matched with all corresponding image materials.
[0075] S104: Display the target image in the image material that matches the image to be queried according to the matching result.
[0076] Optionally, after obtaining the target image that matches the image to be queried, the target image can also be output. When the target image is a video frame, the timestamp and position information of the video frame can also be displayed. Moreover, a folder for saving the target image can be generated, and the target images can be saved in this folder according to the time sorting between the target images, the sorting of the similarity to the image to be queried, or other sorting information.
[0077] Optionally, after displaying the matched target image, the target image can also be provided to the user for processing, and operations such as export, download, content editing, sending, and others can be performed on the target image according to the user's processing operations.
[0078] The image query method of this application has the following advantages:
[0079] 1. Customize the shape of the intercepted video screen: Provide a graphic toolbar for users to draw areas of any shape for interception to meet complex monitoring requirements.
[0080] 2. Automatic search and strong reminder of materials: After importing image materials, they can be automatically searched and strongly reminded, improving the search efficiency and accuracy.
[0081] 3. Flexible interception method: Compared with the traditional rectangular interception, a more practical interception method is provided (such as providing a graphic toolbar, and the cropping area can also be determined through a human-computer interaction method).
[0082] 4. Improve the efficiency of information acquisition: Reduce the workload of manual frame-by-frame inspection and enhance the timeliness and reliability of information acquisition.
[0083] 5. Quickly import search materials: The user selects the images to be imported and the image materials to be queried through buttons on the interface.
[0084] 6. Automatically search for materials: Automatically search for target images that match the imported images to be queried in all loaded video frames.
[0085] 7. Strong reminder: If a matching target image is found in a certain video frame, highlight the target image or pop up a prompt window for strong reminder, and be able to automatically save the current target image without having to stare at and check frame by frame manually.
[0086] 8. Combining multiple feature extraction technologies, multi-scale analysis, temporal consistency, and adaptive parameter adjustment, the image matching of this application is improved in terms of robustness, accuracy, real-time performance, and various other aspects.
[0087] In an alternative embodiment, an electronic device is provided, as Figure 5 shown, Figure 5 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between this electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of this electronic device 4000 does not constitute a limitation to the embodiments of this application.
[0088] The processor 4001 can be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the disclosure of this application. The processor 4001 can also be a combination of computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0089] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0090] The memory 4003 can be a ROM (ReadOnlyMemory), or other types of static storage devices that can store static information and instructions, a RAM (RandomAccessMemory), or other types of dynamic storage devices that can store information and instructions. It can also be an EEPROM (Electrically Erasable ProgrammableReadOnlyMemory), a CD-ROM (CompactDiscReadOnlyMemory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited here.
[0091] The memory 4003 is used to store the computer program for implementing the embodiments of the present application and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0092] Among them, the electronic device can be any kind of electronic product that can perform human-computer interaction with an object. For example, a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an Internet Protocol Television (IPTV), a smart wearable device, etc.
[0093] The electronic device may further include a network device and / or an object device. Among them, the network device includes, but is not limited to, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing (CloudComputing).
[0094] The network where the electronic device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.
[0095] An embodiment of the present application provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiment can be implemented.
[0096] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiment can be implemented.
[0097] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than the illustrated or textually described order.
[0098] It should be understood that although the flowchart of the embodiment of the present application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless there is a clear description in this article, in some implementation scenarios of the embodiment of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiment of the present application does not limit this.
[0099] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, adopting other similar implementation means based on the technical idea of the present application also belongs to the protection scope of the embodiments of the present application.
Claims
1. An image query method, characterized in that, Including: Obtain image materials corresponding to the image to be queried, where the image materials include at least one of videos and pictures; Extract traditional features and depth features in the image to be queried and the image materials according to a preset scale, including: constructing an image pyramid corresponding to the image to be queried and the image materials based on the preset scale, where the preset scale includes at least one of object size and image resolution; extracting the traditional features and the depth features at each scale of the image pyramid, and the traditional features include feature points and descriptors of the image; Generate fusion features corresponding to the traditional features and depth features, including: obtaining the current weights corresponding to the traditional features and depth features, and fusing the traditional features and depth features based on the current weights to generate the fusion features. The current weights are generated based on initial weights, the current scene, the current scale, and error feedback. The fusion features of the image to be queried and the fusion features corresponding to the image materials are obtained separately, and the fusion features are obtained at each scale of the image to be queried and the image materials. The error feedback is determined based on the performance results of the traditional features and depth features during image matching; Perform feature matching between the image to be queried and the image materials based on the fusion features and the preset scale, including: determining the correspondence between the fusion features in the image to be queried and the fusion features in the image materials, generating feature pairs according to the correspondence, and calculating the distance between the two fusion features in the feature pairs; obtaining the matching score at the current scale according to the average distance or weighted distance of all feature pairs at the current scale, and obtaining the matching result according to the weights corresponding to the matching scores at different scales; Display the target image in the image materials that matches the image to be queried according to the matching result.
2. The image query method according to claim 1, wherein The obtaining the image materials corresponding to the image to be queried includes: Obtain the image to be queried according to the obtained image cropping operation; If it is determined that the current meets the preset conditions, obtain the image materials corresponding to the image to be queried according to the preset conditions, where the preset conditions include at least one of image material import, receiving a search instruction, and currently loading image materials.
3. The image query method according to claim 2, wherein The obtaining the image to be queried according to the obtained image cropping operation includes: Load a video file and play the video frame corresponding to the video file; Determine the cropping area in the video frame according to the image cropping operation of the cropping object, and crop the video frame based on the cropping area to obtain the image to be queried.
4. The image query method according to claim 1, wherein The performing feature matching between the image to be queried and the image materials based on the fusion features and the preset scale includes: Match the fusion features of the image to be queried with the fusion features of the image materials at each scale corresponding to the preset scale; Obtain the matching result between the image to be queried and the image materials according to the matching result corresponding to each scale.
5. The image query method according to claim 4, wherein The obtaining the matching result between the image to be queried and the image materials according to the matching result corresponding to each scale includes: Obtain the matching data within a preset time period according to the matching result, where the matching data includes the quantity and quality of the matching result; Adjust the extraction parameters corresponding to the traditional features and the deep features based on the matching data, where the extraction parameters include the feature extraction density.
6. The image query method according to claim 4, wherein The image material is a video. After obtaining the matching result between the image to be queried and the image material according to the matching result corresponding to each scale, it includes: Obtain the matching feature points in each frame according to the matching result; Track the motion trajectory of the feature points according to the timing information of the video, and filter the feature points based on the motion trajectory.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1-6.
8. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-6.
Citation Information
Patent Citations
Fast video object segmentation method and device based on pixel and region feature matching
CN112784750A
Image retrieval method based on multi-feature fusion
CN114140657A
Image retrieval method, device and equipment based on local matching and computer medium
CN117874267A
Text recognition method and device and electronic equipment
CN118230339A