Image processing method, computer program product, storage medium and electronic device
By fusing contextual features of image targets in pedestrian search and using neural networks for similarity calculation, the problem of inaccurate similarity caused by occlusion and lighting changes is solved, thereby improving the accuracy of pedestrian search and the performance of image processing tasks.
Patent Information
- Application Number
- CN202210714375.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-06-22
AI Technical Summary
Existing technologies for pedestrian search suffer from inaccurate pedestrian similarity calculations due to factors such as pedestrian occlusion and changes in lighting, which affects search results.
By acquiring the target's own features in the image to be processed and the reference image, and fusing them with the features of other targets, contextual features are generated. A neural network is used for feature fusion and enhancement, and contextual similarity is calculated to improve the accuracy of similarity calculation.
It improves the accuracy of pedestrian search results by considering the contextual information of the target, thereby enhancing the precision of similarity calculation and the performance of image processing tasks.
Smart Images

Figure CN115205639B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically, to an image processing method, a computer program product, a storage medium, and an electronic device. Background Technology
[0002] When performing pedestrian search, the similarity between the pedestrian to be searched in the image to be searched and the labeled pedestrian in each base image is calculated. The calculated similarity is then sorted, and the pedestrian search results are determined based on the sorting results. The higher the accuracy of the pedestrian similarity calculation, the more likely it is to find the pedestrian to be searched in the base image; otherwise, the more likely it is to find the wrong pedestrian.
[0003] In existing technologies, the features of a specified pedestrian are usually extracted first, and then the pedestrian similarity is calculated based on the extracted features. However, due to many factors, such as occlusion between pedestrians and changes in the external lighting environment, the pedestrian similarity calculated in this way is not accurate, which has a negative impact on pedestrian search results. Summary of the Invention
[0004] The purpose of this application is to provide an image processing method, a computer program product, a storage medium, and an electronic device to improve the above-mentioned technical problems.
[0005] To achieve the above objectives, this application provides the following technical solution:
[0006] In a first aspect, embodiments of this application provide an image processing method, comprising: acquiring the intrinsic features of a target in a to-be-processed image and the intrinsic features of a target in a reference image; wherein the target in the to-be-processed image includes a target to be processed, the target in the reference image includes a reference target, and the target to be processed and the reference target are targets for which similarity is to be calculated; fusing the intrinsic features of the target to be processed with the intrinsic features of at least one other target in the to-be-processed image and / or the reference image to obtain contextual features of the target to be processed; and fusing the intrinsic features of the reference target with the intrinsic features of at least one other target in the to-be-processed image and / or the reference image to obtain contextual features of the reference target; calculating the contextual similarity between the target to be processed and the reference target based on the contextual features of the target to be processed and the contextual features of the reference target; and determining the final similarity between the target to be processed and the reference target based on the contextual similarity between the target to be processed and the reference target.
[0007] In the above method, similarity calculation is not performed directly based on the features of the target to be processed and the features of the reference target. Instead, the features of the target to be processed are first fused with the features of at least one other target in the image to be processed and / or the reference image. This means the contextual information of the target to be processed is incorporated into its features, resulting in the contextual features of the target to be processed. Similarly, the contextual information of the reference target is also incorporated into its features, resulting in the contextual features of the reference target. Then, similarity is calculated based on the contextual features of the target to be processed and the contextual features of the reference target. Because the contextual information of the target is fully considered during the calculation, i.e., the environment in which the target exists is taken into account to evaluate the similarity between targets, the final target similarity score has high accuracy. Using this similarity score to perform image processing tasks may also yield good results.
[0008] It should be noted that the above method does not limit the use of the calculated final similarity. For example, it can be used to perform image processing tasks (such as pedestrian search), or it can simply be stored for later use, and so on.
[0009] In one implementation of the first aspect, the contextual features of the target to be processed are characterized as features of the target to be processed obtained when considering the features of the environment surrounding the target.
[0010] In the above implementation, since the features of the surrounding environment of the target are incorporated when determining the context features of the target, the obtained features can more completely describe the characteristics of the target, which is beneficial to improving the accuracy of similarity calculation.
[0011] In one implementation of the first aspect, fusing the self-features of the target to be processed with the self-features of at least one other target in the image to be processed and / or the reference image to obtain the contextual features of the target to be processed includes: performing feature fusion through one of the following four methods to obtain the contextual features of the target to be processed: fusing the self-features of the target to be processed with the self-features of at least one other target in the image to be processed to obtain the intra-image contextual features of the target to be processed; fusing the self-features of the target to be processed with the self-features of at least one other target in the image to be processed to obtain the intra-image contextual features of the target to be processed; enhancing the intra-image contextual features of the target to be processed to obtain the enhanced features of the target to be processed; fusing the self-features of the target to be processed with the self-features of at least one other target in the image .... The intra-image context features of the target to be processed are obtained by fusing the self-features of another target in the reference image. The intra-image context features of the target to be processed are then fused with the self-features of at least one target in the reference image to obtain the cross-image context features of the target to be processed. The self-features of the target to be processed are then fused with the self-features of at least one other target in the reference image to obtain the intra-image context features of the target to be processed. The intra-image context features of the target to be processed are then fused with the self-features of at least one target in the reference image to obtain the cross-image context features of the target to be processed. The cross-image context features of the target to be processed are then enhanced to obtain the enhanced features of the target to be processed. Herein, the intra-image context features, cross-image context features, and enhanced features of the target to be processed all belong to the context features of the target to be processed.
[0012] The above implementation provides four different feature fusion schemes: intra-graph fusion can be performed only, intra-graph fusion and cross-graph fusion can be performed, and after the features are fused, the features can be enhanced to further improve their expressive power and thus improve the accuracy of similarity calculation.
[0013] In one implementation of the first aspect, calculating the context similarity between the target to be processed and the reference target based on the context features of the target to be processed and the context features of the reference target includes: calculating feature similarity based on each type of feature of the target to be processed and the same type of feature of the reference target, to obtain at least one feature similarity; wherein each feature similarity corresponds to a type of feature; and determining the context similarity between the target to be processed and the reference target based on the at least one feature similarity.
[0014] When calculating contextual similarity, each type of contextual feature can be fully utilized to balance the noise that may be introduced by different features and improve the accuracy of similarity calculation.
[0015] In one implementation of the first aspect, determining the final similarity between the target to be processed and the reference target based on the contextual similarity between the target to be processed and the reference target includes: calculating the self-feature similarity between the target to be processed and the reference target based on the self-feature of the target to be processed and the self-feature of the reference target; calculating the comprehensive similarity between the target to be processed and the reference target based on the contextual similarity and the self-feature similarity; and determining the final similarity between the target to be processed and the reference target based on the comprehensive similarity.
[0016] In the above implementation, the process of calculating the comprehensive similarity based on contextual similarity and original similarity is equivalent to using contextual information to correct the original similarity, which helps to improve the accuracy of similarity calculation.
[0017] In one implementation of the first aspect, determining the final similarity between the target to be processed and the reference target based on the comprehensive similarity between the target to be processed and the reference target includes: obtaining the comprehensive similarity between the target to be processed and the reference target, and the comprehensive similarity between the target to be processed and other targets in the reference image; if the comprehensive similarity between the target to be processed and the reference target is the maximum value among the comprehensive similarities between the target to be processed and all targets in the reference image, then increasing the comprehensive similarity between the target to be processed and the reference target to obtain the final similarity between the target to be processed and the reference target, or directly determining the comprehensive similarity between the target to be processed and the reference target as the final similarity between the target to be processed and the reference target; if the comprehensive similarity between the target to be processed and the reference target is not the maximum value among the comprehensive similarities between the target to be processed and all targets in the reference image, then decreasing the comprehensive similarity between the target to be processed and the reference target to obtain the final similarity between the target to be processed and the reference target.
[0018] Given a reference image, a target to be processed can match at most one target (regardless of whether that target is a reference target). If other targets exist in the reference image, these targets will necessarily not match the target to be processed. The similarity adjustment in the above implementation helps to distinguish the similarity between the target to be processed and the matched target, and the similarity between the target to be processed and the remaining targets, thus facilitating the execution of certain image processing tasks.
[0019] In one implementation of the first aspect, reducing the overall similarity between the target to be processed and the reference target to obtain the final similarity between the target to be processed and the reference target includes: calculating a reduction coefficient corresponding to the reference target based on the overall similarity between the target to be processed and all targets in the reference image; wherein the reduction coefficient is less than 1 and is positively correlated with the overall similarity between the target to be processed and the reference target; and multiplying the overall similarity between the target to be processed and the reference target by the reduction coefficient corresponding to the reference target to obtain the final similarity between the target to be processed and the reference target.
[0020] The above implementation method provides a specific scheme for similarity adjustment.
[0021] In one implementation of the first aspect, fusing the self-features of the target to be processed with the self-features of at least one other target in the image to be processed and / or the reference image to obtain the contextual features of the target to be processed includes: using a feature fusion network to fuse the self-features of the target to be processed with the self-features of at least one other target in the image to be processed and / or the reference image to obtain the contextual features of the target to be processed; wherein the feature fusion network is a neural network, and the structure of the feature fusion network includes one of the following four types: intra-graph feature fusion unit; intra-graph feature fusion unit and feature enhancement unit connected in sequence; intra-graph feature fusion unit and cross-graph feature fusion unit connected in sequence; intra-graph feature fusion unit, cross-graph feature fusion unit and feature enhancement unit connected in sequence.
[0022] The above implementation utilizes a feature fusion network for feature fusion, making the similarity calculation process less reliant on manually set rules and giving it self-learning capabilities, thus improving the accuracy of similarity calculation. The implementation also provides network implementations for four different feature fusion schemes.
[0023] In one implementation of the first aspect, the intra-image context feature fusion unit is used to perform weighted fusion of the self-features of the target to be processed with the self-features of at least one other target in the image to be processed to obtain the intra-image context features of the target to be processed; wherein, the weights used by the self-features in the image to be processed during the weighted fusion in the intra-image feature fusion unit are calculated by the intra-image feature fusion unit based on an attention mechanism according to all features undergoing weighted fusion; the cross-image feature fusion unit is used to perform weighted fusion of the intra-image context features of the target to be processed with the self-features of at least one target in the reference image to obtain the cross-image context features of the target to be processed; wherein, the weights used by the self-features in the reference image during the weighted fusion in the cross-image feature fusion unit are calculated by the cross-image feature fusion unit based on an attention mechanism according to all features undergoing weighted fusion.
[0024] The above implementation introduces an attention mechanism into the intra-graph feature fusion unit and the cross-graph feature fusion unit. The attention mechanism can automatically assign larger weights to important features and smaller weights to unimportant features during feature fusion, thereby enabling the inclusion of truly valuable contextual information in the contextual features, which in turn helps to improve the accuracy of similarity calculation.
[0025] In one implementation of the first aspect, fusing the intrinsic features of the target to be processed with the intrinsic features of at least one other target in the image to be processed and / or the reference image to obtain the contextual features of the target to be processed includes: using a feature fusion network to fuse the intrinsic features of the target to be processed with the intrinsic features of at least one other target in the image to be processed and / or the reference image to obtain the contextual features of the target to be processed; wherein the feature fusion network is a neural network, and the feature fusion network includes at least one feature fusion module connected in sequence, wherein the first feature fusion module is used to fusing the intrinsic features of the target in the image to be processed and the intrinsic features of the target in the reference image based on the intrinsic features of the target input to the module and the intrinsic features of the target in the reference image. Each feature fusion module performs feature fusion based on the features of the target in the image to be processed and the target in the reference image, respectively. Each feature fusion module, except the first one, performs feature fusion based on the features of the target in the image to be processed and the target in the reference image, respectively, to obtain the features of the target in the image to be processed and the target in the reference image, respectively. The structure of the feature fusion module includes one of the following four types: an intra-image feature fusion unit; an intra-image feature fusion unit and a feature enhancement unit connected in sequence; an intra-image feature fusion unit and a cross-image feature fusion unit connected in sequence; or an intra-image feature fusion unit, a cross-image feature fusion unit, and a feature enhancement unit connected in sequence.
[0026] The above implementation utilizes a feature fusion network for feature fusion, making the similarity calculation process less reliant on manually set rules and giving it a self-learning characteristic, thus improving the accuracy of similarity calculation. Furthermore, the implementation uses stacked feature fusion modules for deep feature fusion, which helps to fully incorporate contextual information and further improve the accuracy of similarity calculation.
[0027] In one implementation of the first aspect, the self-features of the target in the image to be processed and the self-features of the target in the reference image are both calculated by a target detection network and a feature extraction network, wherein the target detection network and the feature extraction network are both neural networks. The method further includes: training the target detection network, the feature extraction network, and the feature fusion network; wherein, in the first round of training, the self-features of the target in all training images are calculated using the target detection network and the feature extraction network, and the calculated self-features are written into a cache; in each subsequent round of training, for each training sample pair consisting of two training images, the self-features of the target in all training images are calculated using the target detection network. The network and the feature extraction network calculate the target's own features in the first training image. Using the newly calculated target's own features in the first training image, the target's own features in the first training image already written in the cache are updated. The target's own features in the second training image are directly read from the cache. In the training sample pair, the first training image corresponds to the image to be processed, and the second training image corresponds to the reference image. Both the first training image and the second training image contain multiple targets, and the second training image in each training sample pair is used as the first training image in at least one other training sample pair.
[0028] In the above implementation, the caching mechanism reduces some of the object detection and feature extraction operations, thus accelerating the training process.
[0029] In one implementation of the first aspect, the loss to be calculated during training includes the following three items: target detection loss calculated based on the target detection results of the target detection network; wherein the target detection loss represents the difference between the target detection results and the true target detection results; a first target classification loss calculated based on the target's own features extracted by the feature extraction network; wherein the first target classification loss represents the difference between the target category obtained based on the target's own features and the true target category; a second target classification loss calculated based on the target's context features fused by the feature fusion network; wherein the second target classification loss represents the difference between the target category obtained based on the target's context features and the true target category; or, including the above three items plus: target identity loss calculated based on the similarity of targets in the first training image and the second training image in the same training sample pair; wherein the target identity loss represents the difference between the target identity obtained based on the target similarity and the true target identity, and the target identity refers to whether the target in the first training image and the target in the second training image are the same target.
[0030] In the above implementation, by setting various different losses, effective supervision of the similarity calculation model (including object detection network, feature extraction network, and feature fusion network) is achieved, which is beneficial to improving the performance of the model.
[0031] In one implementation of the first aspect, the method further includes: performing an image processing task using the final similarity between the target to be processed and the reference target.
[0032] In the above implementation, since the calculated final similarity has high accuracy, using this final similarity to perform image processing tasks may yield better results.
[0033] In one implementation of the first aspect, the image processing task is a target search task, the image to be processed is a search image, the target to be processed is a search target, the reference image is a base database image, and the reference target is a labeled target; the step of performing the image processing task using the final similarity between the target to be processed and the reference target includes: sorting the final similarity between the target to be searched and the labeled target calculated based on the image to be searched and each base database image, and determining the target search result based on the sorting result.
[0034] The above implementation provides specific settings for image processing tasks that are target search tasks (e.g., pedestrian search). Since the contextual information of the target (e.g., information of fellow pedestrians) is fully considered when calculating similarity, it is beneficial to improve the search results.
[0035] Secondly, embodiments of this application provide an image processing apparatus, comprising: a feature acquisition component, configured to acquire the self-features of a target in a to-be-processed image and the self-features of a target in a reference image; wherein the target in the to-be-processed image includes a target to be processed, the target in the reference image includes a reference target, and the target to be processed and the reference target are targets for which similarity is to be calculated; a feature fusion component, configured to fuse the self-features of the target to be processed with the self-features of at least one other target in the to-be-processed image and / or the reference image to obtain contextual features of the target to be processed, and to fuse the self-features of the reference target with the self-features of at least one other target in the to-be-processed image and / or the reference image to obtain contextual features of the reference target; a contextual image processing component, configured to calculate the contextual similarity between the target to be processed and the reference target based on the contextual features of the target to be processed and the contextual features of the reference target; and a final image processing component, configured to determine the final similarity between the target to be processed and the reference target based on the contextual similarity between the target to be processed and the reference target.
[0036] Thirdly, embodiments of this application provide a computer program product, including computer program instructions, which, when read and executed by a processor, perform the method provided in the first aspect or any possible implementation of the first aspect.
[0037] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when read and executed by a processor, perform the method provided in the first aspect or any possible implementation thereof.
[0038] Fifthly, embodiments of this application provide an electronic device, including: a memory and a processor, wherein the memory stores computer program instructions, and the computer program instructions are read and executed by the processor to perform the method provided in the first aspect or any possible implementation of the first aspect. Attached Figure Description
[0039] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 The flowchart of an image processing method provided in an embodiment of this application is shown;
[0041] Figure 2 The reasoning process of the similarity calculation model provided in the embodiments of this application is illustrated;
[0042] Figure 3 This illustration shows a possible structure of the feature fusion network provided in an embodiment of this application;
[0043] Figure 4 The training process of the similarity calculation model provided in the embodiments of this application is illustrated;
[0044] Figure 5 The functional modules included in the similarity calculation image processing apparatus provided in the embodiments of this application are shown.
[0045] Figure 6 The possible structure of the electronic device provided in the embodiments of this application is shown. Detailed Implementation
[0046] In recent years, significant progress has been made in research on technologies based on artificial intelligence, such as computer vision, deep learning, machine learning, image processing, and image recognition. Artificial intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies, and application systems to simulate and extend human intelligence. AI is a comprehensive discipline involving numerous technologies, including chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, and neural networks. Computer vision, as an important branch of AI, specifically enables machines to recognize the world. Computer vision technologies typically include face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, object detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, behavior recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, and robot navigation and localization. With the research and advancement of artificial intelligence technology, this technology has been applied in numerous fields, such as security, urban management, traffic management, building management, park management, facial recognition access control, facial recognition attendance, logistics management, warehouse management, robotics, intelligent marketing, computational photography, mobile imaging, cloud services, smart homes, wearable devices, autonomous driving, autonomous driving, smart healthcare, facial payment, facial unlocking, fingerprint unlocking, identity verification, smart screens, smart TVs, cameras, mobile internet, live streaming, beautification, makeup, medical aesthetics, and intelligent temperature measurement. The image processing method in this application embodiment also falls under the category of image processing.
[0047] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. It should be noted that similar reference numerals and letters in the following drawings indicate similar items; therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0048] The terms “comprising,” “including,” or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0049] Figure 1 The flowchart of an image processing method provided in an embodiment of this application is shown. The method includes steps S110 to S150, wherein steps S110 to S140 are used to perform similarity calculation, and step S150 uses the similarity obtained in step S140 to perform an image processing task. This method can be, but is not limited to, [the method described above]. Figure 6 The electronic device in the process performs the operation; for possible structures of this electronic device, please refer to the following section. Figure 6 The explanation.
[0050] It should be noted that the similarity calculated in step S140 does not necessarily have to be used to perform the image processing task in step S150. For example, it could simply be used to store the calculated similarity or for display, etc. In other words, step S150 is merely an optional step that is executed under certain implementations. It is only when the calculated similarity is used to perform the image processing task that the advantages of the image processing method provided in this application's embodiments are more easily illustrated; therefore, steps S110 to S150 are combined and described here. Furthermore, steps S110 to S140 and step S150 may not necessarily be executed on the same electronic device.
[0051] Reference Figure 1 Image processing methods include:
[0052] Step S110: Obtain the target's own features in the image to be processed and the target's own features in the reference image.
[0053] Here, the target refers to a specific object in the image (e.g., pedestrians, animals, vehicles, etc.). The targets in step S110 may all be of the same type (e.g., all pedestrians) or may include multiple types of targets (e.g., some pedestrians, some vehicles). However, the following discussion primarily focuses on the case where all targets are of the same type. Unless otherwise specified, targets mentioned below should be understood as those whose own features were acquired in step S110. If a target, for example, is too small and has been ignored in step S110 without its own features being acquired, then this target need not be considered in subsequent similarity calculations. The target's own features can refer to image features extracted from the region occupied by the target in the image. These features depend on the target's own characteristics and / or its state in the image.
[0054] The target in the image to be processed includes the target to be processed, and the target in the reference image includes the reference target. The target to be processed and the reference target are the targets whose similarity is to be calculated. The number of targets to be processed and reference targets is not limited; there may be one or more. However, the similarity calculation method for each pair of targets (one target to be processed and one reference target) is similar. Therefore, the following text mainly takes the case of only one target to be processed and one reference target as an example.
[0055] The concepts of image to be processed, target to be processed, reference image, and reference target are explained below through two specific image processing tasks. However, it should be understood that image processing tasks are not limited to the following two tasks:
[0056] If the image processing task is a target search task, the image to be processed is the search image, the target to be processed is the search target, and the search image may also contain other targets. Reference images are the base database images; multiple base database images and their necessary supplementary information constitute the base database, which is the scope of the target search. Reference targets are the labeled targets in the base database images (e.g., targets of interest). The base database images may also contain other targets. The possible objectives of a target search task include finding one or more labeled targets in the base database images that are most similar to the search target. This objective can be achieved by calculating the similarity between the search target and the labeled targets.
[0057] If the image processing task is a target tracking task, the image to be processed is the current frame in the video, the target to be processed is all targets in the current frame, the reference image is the previous frame, and the reference target is all targets in the previous frame. The possible objectives of a target tracking task include: determining the correspondence between targets in the current frame and targets in the previous frame, thereby forming the trajectories of different targets. This objective can be achieved by calculating the similarity (pairwise similarity) between targets in the current frame and targets in the previous frame.
[0058] In step S110, there are several ways to obtain the characteristics of the target (including the target to be processed and the reference target):
[0059] For example, features can be extracted from the image during step S110. This could involve first using an object detection network (which could be a neural network) to detect objects in the image, and then using a feature extraction network (which could also be a neural network) to extract features from each detected object. Alternatively, the network model can be replaced with a traditional image feature extraction algorithm.
[0060] For example, features can be extracted before step S110 (which can also be done using object detection networks and feature extraction networks). In step S110, the previously extracted features can be read directly. For instance, in an object search task, the features of the labeled objects in each base image can be extracted in advance. Each time an object search is performed based on the base image, these features can be used directly to calculate similarity without repeating feature extraction.
[0061] Furthermore, in step S110, the image to be processed and the reference image can be acquired first, and then the target's own features in the image to be processed and the target's own features in the reference image can be acquired. However, the step of acquiring the image to be processed and the reference image first is not necessary, because the subsequent similarity calculation only needs to use the target's own features and does not need to use the images themselves.
[0062] For ease of subsequent explanation, the inherent characteristics of the targets can be mathematically represented: Assume there are m targets in the image to be processed and n targets in the reference image, where m and n are both positive integers. The inherent characteristics of all targets in the image to be processed can be represented by the set {pi}, where i takes integers from 1 to m, and pi represents the inherent characteristics of the i-th target in the image to be processed. The inherent characteristics of all targets in the reference image can be represented by the set {qj}, where j takes integers from 1 to n, and qj represents the inherent characteristics of the j-th target in the reference image.
[0063] Step S120: The self-features of the target to be processed are fused with the self-features of at least one other target in the image to be processed and / or the reference image to obtain the contextual features of the target to be processed; and the self-features of the reference target are fused with the self-features of at least one other target in the image to be processed and / or the reference image to obtain the contextual features of the reference target.
[0064] For the target to be processed, "the intrinsic feature of at least one other target in the image to be processed and / or the reference image" refers to the intrinsic feature of at least one other target selected from the intrinsic features of all targets obtained in step S110, excluding the intrinsic features of the target to be processed. The intrinsic features of the selected other target may originate from the image to be processed and / or the reference image.
[0065] The selection of the intrinsic features for fusion is not limited: for example, the intrinsic features of all other targets obtained in step S110 can be used to fuse with the intrinsic features of the target to be processed; for another example, the intrinsic features of other targets in the image to be processed (if there are other targets in the image to be processed) obtained in step S110 can be used to fuse with the intrinsic features of the target to be processed; for yet another example, the intrinsic features of all targets in the reference image obtained in step S110 can be used to fuse with the intrinsic features of the target to be processed; for yet another example, the intrinsic features of those targets (which may originate from the image to be processed or from the reference image) that are close to the target to be processed among all other targets obtained in step S110 can be used to fuse with the intrinsic features of the target to be processed, and so on.
[0066] For the reference target, "the intrinsic features of at least one other target in the image to be processed and / or the reference image" refers to at least one intrinsic feature selected from all the intrinsic features of the targets obtained in step S110, excluding the intrinsic features of the reference target itself. The selected intrinsic features may originate from the image to be processed and / or the reference image. As for how to select the intrinsic features used for fusion, the target to be processed can be referenced.
[0067] It should be noted that the methods for selecting the self-features used for fusion for the target to be processed and the reference target can be the same or corresponding to each other, in order to simplify the feature fusion process (such as simplifying network design): for example, the self-features of other targets in the image to be processed obtained in step S110 can be used to fuse with the self-features of the target to be processed, and the self-features of other targets in the reference image obtained in step S110 can be used to fuse with the self-features of the reference target, but this requirement is not mandatory.
[0068] There are no restrictions on how the various features can be fused. For example, for the target to be processed, the features of the target to be processed and the features of at least one other target obtained in step S110 can be weighted and summed. Alternatively, the features of the target to be processed and the features of at least one other target obtained in step S110 can be fused using a neural network. Examples of feature fusion using neural networks will be given later.
[0069] For the target to be processed, both the features of other targets in the image to be processed and the features of the target in the reference image belong to their corresponding contextual information (the former is contextual information within the same image, and the latter is contextual information across images). This contextual information describes the environment in which the target is located. The feature fusion process in step S120 is equivalent to integrating this environmental description information into the features of the target to be processed, thus the resulting contextual features of the target to be processed have strong expressive power. Similarly, the contextual features of the reference target also incorporate the contextual information corresponding to the reference target, therefore the contextual features of the reference target also have strong expressive power. The stronger the expressive power of the features, the higher the accuracy of the subsequently calculated similarity.
[0070] For example, in one implementation, the contextual features of the target to be processed can be defined as: the features of the target to be processed obtained when considering the features of the surrounding environment. Here, the "features of the surrounding environment" can be reflected by the features of the targets around the target to be processed. These "surrounding targets" can be targets in the image to be processed or targets in the reference image. Different definitions can also be used for "surrounding targets": for example, targets in the image to be processed (or the reference image) that are within a preset distance from the target to be processed are considered targets around the target; another example is targets in the image to be processed (or the reference image) whose image regions overlap with the target to be processed; yet another example is considering all targets in both the image to be processed and the reference image as targets around the target, since images can only capture targets within a certain geographical range, and so on. The contextual features of the reference target can be defined similarly and will not be repeated. Clearly, the contextual features of the target defined in this way incorporate the target's contextual information; however, other methods of defining the contextual features of the target are not excluded.
[0071] The following example illustrates how to calculate the contextual features of the target to be processed:
[0072] Method 1: Fuse the features of the target to be processed with the features of at least one other target in the image to be processed to obtain the intra-image context features of the target to be processed.
[0073] Method 2: First, the features of the target to be processed are fused with the features of at least one other target in the image to obtain the intra-image context features of the target to be processed; then, the intra-image context features of the target to be processed are enhanced to obtain the enhanced features of the target to be processed. Method 2 can also be regarded as adding feature enhancement operations to Method 1.
[0074] Method 3: First, fuse the features of the target to be processed with the features of at least one other target in the image to be processed to obtain the intra-image context features of the target to be processed; then, fuse the intra-image context features of the target to be processed with the features of at least one target in the reference image to obtain the cross-image context features of the target to be processed.
[0075] Method 4: First, the features of the target to be processed are fused with the features of at least one other target in the image to obtain the intra-image context features of the target to be processed. Then, the intra-image context features of the target to be processed are fused with the features of at least one target in the reference image to obtain the cross-image context features of the target to be processed. Finally, the cross-image context features of the target to be processed are enhanced to obtain the enhanced features of the target to be processed. Method 4 can also be regarded as adding feature enhancement operations to Method 3.
[0076] Among them, the intra-image context features, cross-image context features, and enhanced features of the target to be processed all belong to the context features of the target to be processed. Therefore, it is only necessary to execute any one of the above methods to calculate the context features of the target to be processed. It is easy to see that when calculating the intra-image context features of the target to be processed, the features fused all come from the image to be processed, so it is called "intra-image". When calculating the cross-image context features of the target to be processed, the features fused (except for the target itself) all come from the reference image, so it is called "cross-image". Feature enhancement refers to an operation that further extracts features based on input features to enhance the feature representation ability.
[0077] In Method 1, "at least one other target in the image to be processed" can refer to all other targets in the image to be processed, or it can refer to other targets in the image to be processed that meet certain conditions, such as targets that are close to the target to be processed, etc. The "at least one other target" mentioned in Methods 2 to 4 can be interpreted similarly.
[0078] The feature fusion methods in methods 1 to 4 (whether it is intra-graph context feature fusion or cross-graph context feature fusion) can be based on preset rules (e.g., weighted summation based on preset weights), or they can be based on neural networks, etc. Examples of feature fusion using neural networks will be given later.
[0079] In comparison, method 1 is the simplest to implement, while method 4 is the most complex. However, the contextual features of the target obtained in method 4 have the strongest expressive power because they incorporate contextual information within the graph, contextual information across the graph, and feature enhancement.
[0080] The contextual features of the reference target can be calculated by referring to the contextual features of the target to be processed. For example, corresponding to method 4, the following calculation method can be adopted: First, the self-features of the reference target are fused with the self-features of at least one other target in the reference image to obtain the intra-image contextual features of the reference target; then, the intra-image contextual features of the reference target are fused with the self-features of at least one target in the image to be processed to obtain the cross-image contextual features of the reference target; finally, the cross-image contextual features of the reference target are enhanced to obtain the enhanced features of the reference target. Among them, the intra-image contextual features, cross-image contextual features, and enhanced features of the reference target all belong to the contextual features of the reference target.
[0081] The contextual features of a target can be mathematically represented using sets. To represent the intra-graph context features of all targets in the image to be processed, a set is used. To represent the cross-features of all targets in the image to be processed, using a set Let represent the enhanced features of all targets in the image to be processed, where i takes integer values from 1 to m; use a set To represent the intra-graph contextual features of all targets in the reference image, using a set To represent the cross-features of all targets in the reference image, using a set Let j represent the enhanced features of all targets in the reference image, where j takes integer values from 1 to n.
[0082] Regarding the mathematical representation above, it should be noted that:
[0083] First, in step S120, theoretically only the context features of the target to be processed and the context features of the reference target need to be calculated. It is not necessary to calculate the context features of all targets in the image to be processed and the context features of all targets in the reference image. That is, the mathematical representation above may contain some features that do not need to be calculated.
[0084] Second, it is not excluded that some implementations of step S120 will calculate the context features of all targets in the image to be processed and the context features of all targets in the reference image. Even if some targets are neither the target to be processed nor the reference target (for example, the case of the stacked feature fusion module mentioned later), the calculation of the context features of the targets in the image to be processed can refer to the calculation of the context features of the targets to be processed, and the calculation of the context features of the targets in the reference image can refer to the calculation of the context features of the reference targets. This will not be elaborated further.
[0085] Step S130: Calculate the context similarity between the target to be processed and the reference target based on the context features of the target to be processed and the context features of the reference target.
[0086] If features are represented in the form of feature vectors, the similarity between two features can be expressed as the inner product of the two feature vectors, the cosine value, the distance value, etc., or even by using a dedicated neural network to calculate the similarity. The following text mainly uses the inner product form as an example.
[0087] In one implementation, if the context features of the target to be processed include at least one type of feature (for example, method 1 includes one type of feature, namely the intra-graph context features of the target to be processed, and method 4 includes three types of features, namely the intra-graph context features of the target to be processed, cross-graph context features, and augmentation features), and the types of features included in the context features of the target to be processed are the same as the types of features included in the context features of the reference target, then the context similarity between the target to be processed and the reference target can be calculated as follows:
[0088] First, based on each type of feature of the target to be processed and the same type of feature of the reference target, the feature similarity between the target to be processed and the reference target for that type of feature is calculated, and at least one feature similarity is obtained (there are at least one type of feature, and one feature similarity is calculated for each type of feature, so there is at least one feature similarity); then, the context similarity between the target to be processed and the reference target is determined based on the at least one feature similarity.
[0089] For example, the contextual features of the target to be processed include in-graph contextual features. Cross-graph context features and enhanced features The contextual features of the reference target include in-graph contextual features. Cross-graph context features and enhanced features First, it can be based on and Calculate the intra-graph contextual feature similarity between the target to be processed and the reference target. (Inner product), according to and Calculate the cross-graph contextual feature similarity between the target to be processed and the reference target. according to and Calculate the enhanced feature similarity between the target to be processed and the reference target. Then, the mean of the above three similarities is calculated to obtain the contextual similarity between the target to be processed and the reference target. Of course, the mean can be replaced with other operations.
[0090] Note that in this example, the i-th target in the image to be processed is taken as the target to be processed, and the j-th target in the reference image is taken as the reference target. However, since i can actually take any integer from 1 to m, and j can also take any integer from 1 to n, the context similarity can actually be calculated in the above manner for any target in the image to be processed and any target in the reference image, regardless of whether these two targets are the target to be processed and the reference target. In step S130, theoretically only the context similarity between the target to be processed and the reference target needs to be calculated (but it is not excluded that the context similarity between other targets will be calculated in some implementations).
[0091] The above methods for calculating context similarity make full use of each type of context feature, thereby balancing the noise that may be introduced from different features and improving the accuracy of subsequent similarity calculations.
[0092] It should be understood that there are other methods for similarity calculation. For example, using the example above, we can... and Merging into one feature Will and It also merges into a single feature. Then according to and Calculate the contextual similarity between the target to be processed and the reference target.
[0093] Step S140: Determine the final similarity between the target to be processed and the reference target based on the contextual similarity between the target to be processed and the reference target.
[0094] If we denote the contextual similarity between the target to be processed and the reference target as... The final similarity between the target to be processed and the reference target is denoted as . One way to implement step S140 is to directly use the contextual similarity between the target to be processed and the reference target as the final similarity between the target to be processed and the reference target, that is... Or It requires some calculations to obtain But no matter how you get it Because in calculation It fully considers the contextual information of the target, that is, it combines the environment in which the target is located to assess the similarity between targets, so It has high accuracy (able to accurately assess the similarity between the target to be processed and the reference target), thus enabling subsequent utilization. Performing image processing tasks may also yield good results.
[0095] The following describes a possible method. calculate Way:
[0096] Step A: Calculate the similarity between the self-features of the target to be processed and the self-features of the reference target.
[0097] For example, the characteristic of the target to be processed is p. i The reference target's own characteristics are q. j Then the similarity of the self-features between the target to be processed and the reference target can be Obviously, The text does not include contextual information.
[0098] Step B: Calculate the comprehensive similarity between the target to be processed and the reference target based on the contextual similarity and self-feature similarity between the target to be processed and the reference target.
[0099] For example, if the comprehensive similarity is calculated by weighting contextual similarity and self-feature similarity, then step B can be expressed by the formula: Where s i,j λ represents the overall similarity between the target to be processed and the reference target, where λ is the weighted calculation weight, which can take values in the range [0,1], such as 0.5, 0.6, etc. i,j The calculation process can also be viewed as utilizing contextual information. For those who do not consider context information The process of adjusting the values. It should be understood that methods other than weighted methods can also be used to calculate the overall similarity.
[0100] Step C: Determine the final similarity between the target to be processed and the reference target based on the comprehensive similarity between the target to be processed and the reference target.
[0101] One way to implement step C is to directly use the combined similarity between the target to be processed and the reference target as the final similarity between the target to be processed and the reference target, that is... Or s i,j It requires some calculations to obtain But no matter how you get it Because in calculating s i,j At the same time, and This approach considers both contextual information and the inherent characteristics themselves, preventing an overemphasis on contextual information and thus improving performance. The calculation accuracy will be given later based on s. i,j calculate Examples will not be elaborated here.
[0102] It should be understood that, apart from steps A through C, according to calculate There are other ways, for example, to Multiply by the adjustment factor to get etc.
[0103] Step S150: Perform image processing tasks using the final similarity between the target to be processed and the reference target.
[0104] The specific method for utilizing the final similarity of targets to perform image processing tasks depends on the specific content of the image processing task. For example, in a target search task, where the image to be processed is the image to be searched, the target to be processed is the target to be searched, the reference images are the base images, and the reference targets are the labeled targets, then in step S150, the final similarities calculated between the image to be searched and each base image, and between the target to be searched and the labeled targets, can be sorted, and the target search results can be determined based on the sorting results. For example, if there are 1000 base images, and assuming that each base image contains only one labeled target, then a total of 1000 final similarities will be calculated. These final similarities can be sorted in descending order, and the largest one or several can be taken as the search results, meaning that these labeled targets and the target to be searched are the same target.
[0105] In summary, Figure 1 The image processing method in this paper considers the contextual information of the target in the process of calculating the final similarity, so the calculated final similarity has high accuracy. Therefore, using the final similarity to perform image processing tasks may yield better results.
[0106] For example, in pedestrian search tasks, if we directly calculate the similarity of the self-features of the pedestrian to be searched and the self-features of the labeled pedestrian, and then use this similarity for the search, the ability to represent the self-features may be poor if there is occlusion between the pedestrians. Therefore, the self-feature similarity may not effectively describe the true similarity between the pedestrian to be searched and the labeled pedestrian. Figure 1 In this method, the context features of the pedestrian to be searched and the labeled pedestrian are used when calculating the final similarity between the pedestrian to be searched and the labeled pedestrian. The context features contain the context information of the pedestrian, or the information of the pedestrians (including people around the pedestrian to be searched and people around the labeled pedestrian). Therefore, its expressive power is significantly stronger than the pedestrian's own features, and the accuracy of the final similarity calculated is also higher, which is beneficial to improving the pedestrian search results.
[0107] Below, based on the above embodiments, a method of step C described above, consisting of s i,j calculate Way:
[0108] Step C1: Obtain the overall similarity between the target to be processed and each target in the reference image.
[0109] Although only the comprehensive similarity s between the target to be processed and the reference target was calculated in step B. i,j However, since i and j can actually take any value, the comprehensive similarity between any target in the image to be processed and any target in the reference image can be calculated in a similar way. Thus, the comprehensive similarity between the target to be processed and each target in the reference image can also be calculated (a total of n comprehensive similarities will be calculated).
[0110] Step C2: If the overall similarity between the target to be processed and the reference target is the maximum value among the overall similarities between the target to be processed and all targets in the reference image, then increase the overall similarity between the target to be processed and the reference target to obtain the final similarity between the target to be processed and the reference target; or, directly determine the overall similarity between the target to be processed and the reference target as the final similarity between the target to be processed and the reference target. If the overall similarity between the target to be processed and the reference target is not the maximum value among the overall similarities between the target to be processed and all targets in the reference image, then decrease the overall similarity between the target to be processed and the reference target to obtain the final similarity between the target to be processed and the reference target.
[0111] To increase the overall similarity between the target and the reference target, you can multiply the overall similarity by an increase factor greater than 1, or you can add an adjustment value greater than 0 to the overall similarity, or you can calculate the power of the overall similarity, and so on. To decrease the overall similarity between the target and the reference target, you can multiply the overall similarity by a decrease factor less than 1, or you can add an adjustment value less than 0 to the overall similarity, or you can calculate the power of the overall similarity, and so on.
[0112] The following section provides a brief analysis of the target search task for s. i,j The reason for adjusting the value is:
[0113] To simplify the description, let's call the target with the highest overall similarity in the reference image the matched target, and the targets in the reference image other than the matched target the remaining targets. Note that the matched target is not necessarily the reference target.
[0114] For a given reference image, a target to be searched can match at most one previous target (the matched target). If there are other targets in the reference image (the remaining targets), then these targets must not match the target to be processed (for example, the same person cannot appear twice in an image).
[0115] The similarity adjustment described above helps distinguish between the final similarity between the search target and the matching target, and the final similarity between the search target and the remaining targets. Therefore, when ranking the final similarity (assuming descending order), if a reference target belongs to the remaining targets, its final similarity ranking will be lower, making it less likely to be selected as a search result. In other words, the search result is more accurate. Conversely, if the similarity is not adjusted, even if a reference target in a reference image is a remaining target, it might be similar to the search target, meaning its final similarity is also relatively high (though not as high as the matching target). This reference target would still be ranked relatively high and likely to be selected as a search result, but it would clearly be an incorrect search result.
[0116] If the overall similarity between the target to be processed and the reference target (which is the remaining target) is reduced by multiplying by a reduction factor, the reduction factor can be calculated and the similarity adjustment can be completed as follows:
[0117] First, based on the comprehensive similarity between the target to be processed and each target in the reference image, the reduction coefficient corresponding to each remaining target is calculated. Then, the comprehensive similarity between the target to be processed and each remaining target is multiplied by the reduction coefficient corresponding to that remaining target to obtain the final similarity between the target to be processed and each remaining target, which naturally also includes the final similarity between the target to be processed and the reference target (at this time, the reference target is the remaining target).
[0118] The reduction coefficient is less than 1 and is positively correlated with the overall similarity between the target to be processed and the remaining targets. The specific calculation method for the reduction coefficient is not limited; for example, it can be calculated using the following formula:
[0119]
[0120]
[0121] In the first formula above, c i,j It is s i,j The result after processing by the softmax function; the softmax function simply adjusts s i,j Mapping to the interval (0,1), the values of s before and after the mapping i,j The size relationship between them remains unchanged. Σ j In the denominator, j takes integers from 1 to n, meaning that the denominator uses the combined similarity between the target to be processed and each target in the reference image obtained in step C1.
[0122] The max of the second formula above j In the expression, j takes integers from 1 to n, and max... jc i,j This means taking all c. i,j The maximum value in, This represents the reduction factor, which, based on the definition of the matching target and the property that the softmax function does not change the size relationship before and after the mapping, represents the reduction factor c of the matching target. i,j It's max j c i,j Therefore, the reduction factor for the target is 1, meaning that for the target... Alternatively, it can be considered that the matching target does not need to use a reduction factor. For each remaining target, its corresponding c i,j It must be less than the maximum. j c i,j Therefore, its corresponding reduction factor is less than 1, and c i,j The smaller (equivalent to s) i,j The smaller the value, the smaller the reduction factor, s i,j The greater the reduction (i.e., the value of the reduction coefficient is positively correlated with the overall similarity between the target to be processed and the remaining targets).
[0123] It should be noted that, theoretically, only the final similarity between the target to be processed and the reference target needs to be calculated. It is not necessary to calculate the final similarity between the target to be processed and each target in the reference image as in the formula above (because i and j can take any value). The above formula is used only for simplification.
[0124] It should be understood that there are many other ways to calculate the reduction factor. For example, it can be calculated using the following formula (the reduction factor is...). ):
[0125]
[0126] For example, it can be calculated using the following formula (the reduction factor is...). ):
[0127]
[0128] etc.
[0129] Furthermore, it should be understood that in step C, based on s i,j calculate There are other ways, for example, for s i,j Multiply by the adjustment factor to get etc.
[0130] Below, based on the above embodiments, we will continue to introduce a scheme for similarity calculation using neural networks:
[0131] In some implementations, a pre-trained neural network can be used to execute step S120; this can be called a feature fusion network. Using a feature fusion network for feature fusion makes the similarity calculation process less reliant on manually set rules, giving it a self-learning characteristic. This avoids noise and errors introduced by humans, thus improving the accuracy of similarity calculation.
[0132] As mentioned earlier, the acquisition of the target's own features in step S110 can be achieved through an object detection network and a feature extraction network. Therefore, the object detection network, feature extraction network, and feature fusion network can be considered as constituting a larger neural network, which can be called a similarity calculation model. This model can be trained and inferred as a whole. Figure 2 and Figure 4 The reasoning and training processes of the similarity calculation model are shown respectively.
[0133] Reference Figure 2 Ignoring the internal structure of the feature fusion network, firstly, the target detection network and feature extraction network can be used to obtain the self-features of the target in the image to be processed and the self-features of the target in the reference image. Then, the feature fusion network can be used to fuse the self-features of the target to be processed with the self-features of at least one other target in the image to be processed and / or the reference image to obtain the contextual features of the target to be processed. Similarly, the self-features of the reference target can be fused with the self-features of at least one other target in the image to be processed and / or the reference image to obtain the contextual features of the reference target. Finally, the final similarity between the two can be calculated based on the contextual features of the target to be processed and the contextual features of the reference target.
[0134] Furthermore, the structure of the feature fusion network can adopt, but is not limited to, one of the following four:
[0135] (1) Includes in-figure feature fusion units;
[0136] (2) It includes a graph feature fusion unit and a feature enhancement unit connected in sequence;
[0137] (3) It includes intra-graph feature fusion units and cross-graph feature fusion units connected in sequence;
[0138] (4) It includes a graph-intra-graph feature fusion unit, a cross-graph feature fusion unit, and a feature enhancement unit connected in sequence.
[0139] All of the above structures (1) to (4) can perform feature fusion to calculate the context features of the target to be processed and the context features of the reference target. Only one can be chosen when implementing the feature fusion network. The functions of each unit mentioned in structures (1) to (4) in calculating the context features of the target to be processed are as follows:
[0140] Intra-image feature fusion unit: used to fuse the features of the target to be processed with the features of at least one other target in the image to be processed to obtain the intra-image context features of the target to be processed.
[0141] Cross-graph feature fusion unit: used to fuse the intra-graph context features of the target to be processed with the self-feature features of at least one target in the reference image to obtain the cross-graph context features of the target to be processed.
[0142] Feature enhancement unit: Used to enhance the features of the target to be processed output by the previous unit, so as to obtain the enhanced features of the target to be processed. The "previous unit" here refers to the intra-graph feature fusion unit in structure (2) and the cross-graph feature fusion unit in structure (4).
[0143] Among them, the intra-graph context features, cross-graph context features, and enhanced features of the target to be processed all belong to the context features of the target to be processed. It is not difficult to see that the feature fusion effect achieved by structures (1) to (4) corresponds to methods 1 to 4 in the previous text, or in other words, structures (1) to (4) are a specific implementation of methods 1 to 4 using neural networks. Figure 2 and Figure 4 In the process, the feature fusion network adopted structure (3), while Figure 3 In this context, the feature fusion network adopts structure (4).
[0144] The functions of each unit mentioned in structures (1) to (4) in calculating the context features of the reference target are similar to those in calculating the context features of the target to be processed. It is only necessary to interchange the image to be processed, the target to be processed, and the reference image and the reference target in the above functional descriptions, and will not be repeated.
[0145] Continue to refer to Figure 2 In the intra-image feature fusion unit, the left side is the image to be processed, where each circle represents a target. The black circle represents the target to be processed, and the target's own features are fused with the own features (dashed lines) of other targets in the image to be processed to obtain the intra-image context features of the target to be processed. The right side is the reference image, where each circle represents a target, and the gray circle represents the reference target. The target's own features are fused with the own features (dashed lines) of other targets in the reference image to obtain the intra-image context features of the reference target.
[0146] In the cross-graph feature fusion unit, the left side is the image to be processed. The intra-graph context features of the target to be processed are fused with the intra-graph context features (dashed lines) of all targets in the reference image to obtain the cross-graph context features of the target to be processed. The right side is the reference image. The intra-graph context features of the reference target are fused with the intra-graph context features (dashed lines) of all targets in the image to be processed to obtain the cross-graph context features of the reference target.
[0147] The output of the feature fusion network includes the context features of the target to be processed (intra-graph context features and cross-graph context features of the target to be processed), and the context features of the reference target (intra-graph context features and cross-graph context features of the reference target).
[0148] Furthermore, attention mechanisms can be introduced into intra-graph feature fusion units and cross-graph feature fusion units. In this case, the functions of these two units in calculating the contextual features of the target to be processed are as follows:
[0149] Intra-image feature fusion unit: used to perform weighted fusion of the self-features of the target to be processed with the self-features of at least one other target in the image to be processed, to obtain the intra-image context features of the target to be processed; wherein, the weights used by the self-features in the image to be processed in the intra-image feature fusion unit when performing weighted fusion are calculated by the intra-image feature fusion unit based on an attention mechanism according to all features to be weighted and fused.
[0150] Cross-graph feature fusion unit: used to perform weighted fusion of the intra-graph context features of the target to be processed with the self-features of at least one target in the reference image to obtain the cross-graph context features of the target to be processed; wherein, the weights used by the self-features in the reference image in the cross-graph feature fusion unit during weighted fusion are calculated by the cross-graph feature fusion unit based on an attention mechanism according to all features to be weighted (the target to be processed may not calculate the corresponding weights itself).
[0151] Reference Figure 3 Let's understand the attention mechanisms in intra-graph feature fusion units and cross-graph feature fusion units. First, let's look at... Figure 3 The leftmost column shows the in-image feature fusion unit, which includes a multi-head attention structure within a transformer and a summation and layer normalization structure. Its input is the target's own features {p} in the image to be processed. i The output is the intra-image context features of the target in the image to be processed. This naturally includes the intra-image context features of the target to be processed (in principle, the intra-image context features of other targets in the image to be processed can also be omitted, but they are all written here for simplicity; the cross-image context features and enhancement features are similar). The calculation process of the intra-image feature fusion unit can be represented by the following formula:
[0152] e i,j =f Q (p i ) T f K (p j )
[0153]
[0154]
[0155] Since it is an intra-image feature fusion unit, it involves features in the image to be processed, that is, the values of i and j are integers from 1 to m.
[0156] The ∑ in the first, second, and third formulas above j w i,j f V (p j The part corresponding to the operation performed by the multi-head attention structure, f Q and f K These are the functions used by the multi-head attention structure to compute the Q-vector and K-vector, respectively. i,j This represents the similarity (inner product) between vectors Q and K, while w i,j The multi-head attention structure processes e through the softmax function. i,j The final result is the weight corresponding to each feature in the image to be processed. This weight reflects the difference in importance of different features during weighted summation, i.e., attention. V It is the function used by the multi-head attention structure to compute the V vector, ∑ j w i,j f V (p j The expression represents a weighted average of the features of all targets in the image to be processed. It should be noted that multi-head attention structures typically calculate multiple sets of Q / K / V vectors (the so-called "multi-head"), but for simplicity, the formula here only calculates one set of Q / K / V vectors, although this approach is still feasible.
[0157] The remainder of the third formula above corresponds to the summation and layer normalization structure, where LN represents the layer normalization operation.
[0158] The cross-graph feature fusion unit comprises a multi-head attention structure within a transformer and a summation and layer normalization structure. Its input is the intra-graph contextual features of the target in the image to be processed. The intrinsic features of the target in the reference image {q j The output is the cross-graph contextual features of the target in the image to be processed. This naturally includes the cross-graph contextual features of the target object. The calculation process of the cross-graph feature fusion unit can be expressed by the following formula:
[0159] e′ i,j =f Q (p i )T f K (q j )
[0160]
[0161]
[0162] Since it is a cross-image feature fusion unit, it involves features from the image to be processed and the reference image, i.e., the value of i is an integer from 1 to m, and the value of j is an integer from 1 to n.
[0163] The ∑ in the first, second, and third formulas above j w′ i,j f V (q j The part corresponding to the operation performed by the multi-head attention structure, f Q and f K These are the functions used by the multi-head attention structure to compute the Q-vector and K-vector, respectively, e′ i,j This represents the similarity between vector Q and vector K, while w′ i,j The multi-head attention structure processes e′ through the softmax function. i,j The final result, which is the weight corresponding to each feature in the reference image, reflects the difference in importance of different features during weighted summation, i.e., attention. V It is the function used by the multi-head attention structure to compute the V vector, ∑ j w′ i,j f V (q j This indicates that the features of all targets in the reference image are weighted.
[0164] The remainder of the third formula above corresponds to the summation and layer normalization structure, where LN represents the layer normalization operation.
[0165] contrast Figure 3 The intra-image feature fusion unit and the cross-image feature fusion unit in the image have the same network structure, but different inputs. The Q vector, K vector and V vector of the intra-image feature fusion unit are calculated based on the features of the target in the image to be processed, while only the Q vector of the cross-image feature fusion unit is calculated based on the features of the target in the image to be processed, while the K vector and V vector are calculated based on the features of the target in the reference.
[0166] In summary, by introducing the attention mechanism into the intra-graph feature fusion unit and the cross-graph feature fusion unit, the attention mechanism can automatically assign larger weights to important features and smaller weights to unimportant features during feature fusion. This allows truly valuable contextual information to be incorporated into the contextual features, thereby improving the accuracy of similarity calculation.
[0167] In addition, since the attention mechanism can automatically filter out some targets that are not closely related to similarity calculation through weight allocation (for example, targets that are far away from the target to be processed have very small weights, which is equivalent to not considering the target when feature fusion), all targets in the image to be processed can be considered when calculating intra-image context features, and all targets in the reference image can be considered when calculating cross-image context features, without the need for deliberate screening.
[0168] Continue to refer to Figure 3 , Figure 3 The feature fusion network also includes a feature enhancement unit, which comprises a multi-layer perceptron (MLP), a summation and layer normalization structure, and an L2 normalization structure (L2 Norm). Its input is the cross-image context features of the target in the image to be processed. The output is the enhanced features of the target in the image to be processed. This naturally includes the enhanced features of the target object. The calculation process of the feature enhancement unit can be expressed by a formula:
[0169]
[0170] In the formula above, L2 represents the L2 normalization operation, LN represents the layer normalization operation, and MLP represents the computation process of the multilayer perceptron.
[0171] Let's look again. Figure 3 The rightmost column has the same network structure as the leftmost column, except that the feature fusion object is now the features in the reference image; the fusion process will not be described in detail here. It should also be noted that although... Figure 3 The diagram shows two intra-graph feature fusion units (each including a multi-head attention structure and a summation and layer normalization structure), but this is only for ease of understanding. The feature fusion network can implement only one intra-graph feature fusion unit, which can calculate the intra-graph context features of the target in the image to be processed and the intra-graph context features of the target in the reference image, respectively. For Figure 3 The cross-graph feature fusion unit and feature enhancement unit in the graph can be understood in a similar way.
[0172] In addition to the implementation methods described above, the feature fusion network can also adopt a stacked structure, as follows:
[0173] The feature fusion network includes at least one feature fusion module connected in sequence. The first feature fusion module takes the features of the target in the image to be processed and the features of the target in the reference image as input, performs feature fusion, and obtains its own output features of the target in the image to be processed and the target in the reference image. Each feature fusion module other than the first takes the features of the target in the image to be processed and the features of the target in the reference image output by the previous at least one feature fusion module as input, performs feature fusion, and obtains its own output features of the target in the image to be processed and the target in the reference image. Note that the number of features input and output by each feature fusion module is the same, for example, m+n features (corresponding to m+n targets in the image to be processed and the reference image), so these feature fusion modules can be connected together. The structure of the feature fusion module includes one of the following four:
[0174] (a) Intra-graphic feature fusion unit;
[0175] (b) In-graph feature fusion units and feature enhancement units connected in sequence;
[0176] (c) Intra-graph feature fusion units and cross-graph feature fusion units connected sequentially;
[0177] (d) Intra-graph feature fusion unit, cross-graph feature fusion unit and feature enhancement unit connected in sequence.
[0178] Structures (a) to (d) can be the same as structures (1) to (4) mentioned above, but the inputs and outputs may not be the same.
[0179] The functions of each unit mentioned in structures (a) to (d) in calculating the contextual features of the target in the image to be processed are as follows (taking the case where "at least one previous feature fusion module" only contains the previous feature fusion module as an example):
[0180] Intra-image feature fusion unit: used to fuse the features of each target in the image to be processed output by the previous feature fusion module of the current feature fusion module with the features of at least one other target in the image to be processed output by the previous feature fusion module to obtain the intra-image context features of the target in the image to be processed.
[0181] Specifically, if the feature fusion network has only one feature fusion module, the intra-graph feature fusion unit of the feature fusion module is used to fuse the self-feature of each target in the image to be processed with the self-feature of at least one other target in the image to be processed, so as to obtain the intra-graph context features of the target in the image to be processed.
[0182] Cross-graph feature fusion unit: used to fuse the intra-graph context features of each target in the image to be processed with the features of at least one target in the reference image output by the previous feature fusion module of the current feature fusion module to obtain the cross-graph context features of the target in the image to be processed.
[0183] Specifically, if the feature fusion network has only one feature fusion module, then the cross-graph feature fusion unit of that module is used to fuse the intra-graph context features of the target in the image to be processed with the self-feature features of at least one target in the reference image to obtain the cross-graph context features of the target in the image to be processed. The feature enhancement unit is used to enhance the features of the target in the image to be processed output by the previous unit to obtain the enhanced features of the target in the image to be processed. Here, "the previous unit" refers to the intra-graph feature fusion unit in structure (b) and the cross-graph feature fusion unit in structure (d).
[0184] All the intra-graph context features, cross-graph context features, and enhanced features of the target in the image to be processed calculated by the feature fusion modules belong to the context features of the target in the image to be processed, and the context features of the target in the image to be processed include the context features of the target to be processed.
[0185] The functions of each unit mentioned in structures (a) to (d) in calculating the contextual features of the target in the reference image are similar to those in calculating the contextual features of the target in the image to be processed. It is only necessary to interchange the image to be processed and the reference image in the above functional descriptions, and will not be repeated here.
[0186] It is easy to see that if a feature fusion network includes multiple feature fusion modules, and each feature fusion module is implemented using structure (d), then the feature fusion network can be regarded as being composed of... Figure 3 The structure is stacked in the middle ( Figure 3 It can be considered as having only one feature fusion module using structure (d), and, because Figure 3 In fact, it can compute the contextual features of all targets in the image to be processed and the reference image (not just the contextual features of the target to be processed and the reference target), so stacking Figure 3 The structure is not difficult to implement. The same applies to the implementation of each feature fusion module using structures (a), (b), or (c), so the analysis will not be repeated.
[0187] In summary, deep fusion of features through stacked feature fusion modules helps to better integrate contextual information and improve the accuracy of similarity calculation.
[0188] Continue to refer to Figure 2For the image to be processed in the intra-graph feature fusion unit, the four targets can be considered as four nodes. Connecting these four nodes pairwise with edges logically forms a completely undirected graph. The features corresponding to each node (i.e., the target's own features) are fused with the features of adjacent nodes during the fusion process. In other words, the intra-graph feature fusion unit can be viewed as constructing a hidden graph neural network (GNN) for feature fusion. Therefore, as an alternative, a graph neural network can also be explicitly constructed for feature fusion, and an attention mechanism (e.g., a graph convolutional network, GAT) can be added to this network. Similar alternatives exist for cross-graph feature fusion units.
[0189] Below, based on the above embodiments, we will continue to describe the training process of the similarity calculation model. Note that the model training steps and steps S110 to S150 may not be executed on the same electronic device.
[0190] Since the similarity calculation model uses a pair of images (the image to be processed and the reference image) for each inference step, it also needs to be trained using training sample pairs consisting of two images from the training set (referred to as training images or training samples). The first training image in the training sample pair corresponds to the image to be processed, and the second training image corresponds to the reference image. It doesn't matter which training image in the training sample pair is the "first" one. The two images in the training sample pair should ideally contain multiple targets and have many identical targets. This is because if the training images only contain a single target, according to the scheme of this application, the contextual information is very limited (at least there is no intra-image contextual information), and the performance of the trained similarity calculation model will naturally be poor.
[0191] In one possible training scheme, the training of the similarity calculation model is divided into multiple epochs. In each epoch (except for the first epoch), all training sample pairs are traversed, and the model's prediction loss is calculated and the model parameters are updated based on these training sample pairs until training terminates when certain conditions are met (e.g., model convergence).
[0192] The first round of training is unique, as it can be performed directly on a single training image without needing to form the aforementioned training sample pairs (although forming training sample pairs is also possible). During the first round, each training image is processed through the entire similarity calculation model; that is, object detection, feature extraction, and feature fusion are performed on each image, but similarity is not calculated (because similarity cannot be calculated on a single training image). The output at this stage includes the object detection results of the training image, the intrinsic features of the object in the training image, and the contextual features of the object in the training image. Based on these outputs, loss calculations can be performed, and the parameters of the similarity calculation model can be updated. Which losses can be calculated will be explained later.
[0193] The features of the target in the training image are calculated and then written to a cache for later use. Here, "cache" generally refers to a region of storage medium that stores data. Of course, in addition to writing the features themselves to the cache, the identifiers of the training image and the target can also be written to facilitate subsequent feature queries.
[0194] The aforementioned "each training image used in training" refers to each training image that will be used to form training sample pairs in subsequent training rounds.
[0195] Each subsequent training round requires training using the aforementioned training sample pair. For the first training image in the training sample pair, the entire image processing network should be used to perform object detection, feature extraction, and feature fusion on the first training image. The output at this point includes the object detection result of the first training image, the object's intrinsic features in the first training image, and the object's contextual features in the first training image, such as... Figure 4 As shown in the column on the left.
[0196] For the intrinsic features of the target in the first training image, after calculation, the corresponding intrinsic features in the cache are updated. That is, the intrinsic features of the target in the first training image that were written in the previous training round. For example, in the first round of training, the intrinsic features of target X in the first training image are written into the cache. In the second round of training, since the parameters of the similarity calculation model have been updated, the calculated intrinsic features of the same target X in the same training image will naturally be different. The new intrinsic features of X can be used to overwrite the old intrinsic features (of course, the old features can also be discarded if not overwritten). As explained earlier, since the cache can store the identifiers of the training images and the identifiers of the targets, it is easy to retrieve the intrinsic features of X that were previously written into the cache and update them accordingly.
[0197] For the second training image in the training sample pair, instead of using the object detection network and feature extraction network, the target's own features from the second training image, which were written in the previous training round, are directly read from the cache. These features are then fed into the feature fusion network to calculate the contextual features of the target in the second training image, such as... Figure 4 As shown in the column on the right. For example, during the first round of training, the features of target Y in the second training image are written into the cache. During the second round of training, when processing the second training image, the features of Y can be directly read and fused. As explained earlier, since the cache can store the identifiers of the training images and the target, it is easy to retrieve and read the features of Y that were previously written into the cache.
[0198] Furthermore, based on the contextual features of the target in the first training image and the contextual features of the target in the second training image, similarity (such as contextual similarity, comprehensive similarity, and final similarity) can be calculated. Thus, for a training sample pair, the output (excluding those read from the cache) can include the target detection result of the first training image, the target's own features in the first training image, the contextual features of the target in the first training image, the contextual features of the target in the second training image, and the similarity between the targets in the first and second training images (which can be pairwise similarity). Based on these outputs, loss can be calculated and the parameters of the similarity calculation model can be updated. As for which losses can be calculated, this will be explained later.
[0199] Note that it is optional to calculate similarity during training. As can be seen from the following text, even without calculating similarity, model training can still be completed (by calculating loss (1) to (3)).
[0200] Furthermore, it should be noted that if training is performed solely as described above, the target's own features in the second training image of the training sample pair will never be updated. To address this issue, it is required that when constructing training sample pairs, the second training image in each training sample pair should be used as the first training image in at least one other training sample pair, so that the target's own features in each training image will be updated in one round of training.
[0201] In the training method of the aforementioned similarity calculation model, by setting a cache, the operations of object detection and feature extraction for some training images are reduced. Features can be directly read from the cache, thus accelerating the training process. Since the scheme in this application requires acquiring a large number of target features, this training method significantly improves training efficiency. Of course, without setting a cache, object detection and feature extraction can also be performed on both training images in each training sample pair.
[0202] The following is a brief introduction to the loss that can be calculated during training:
[0203] In one implementation, the calculable loss includes the following three items:
[0204] (1) Target detection loss calculated based on the target detection results of the target detection network.
[0205] Among them, the target detection loss characterizes the difference between the target detection result and the true target detection result (which needs to be labeled in advance). The specific form of the loss function is not limited, and this loss can affect the parameter update in the target detection network.
[0206] In the first round of training, each training image corresponds to an object detection result, and the object detection loss can be calculated. In subsequent rounds of training, only the first training image in each training sample pair corresponds to an object detection result, and the object detection loss can be calculated.
[0207] (2) The first target classification loss is calculated based on the target's own features extracted by the feature extraction network.
[0208] The first target classification loss represents the difference between the target category obtained based on the target's own features and the true target category (which needs to be labeled in advance). The specific form of the loss function is not limited. This loss can affect the parameter updates in the target detection network and the feature extraction network.
[0209] For example, each target in the training image can be labeled with an ID, and each ID represents a target category. This is the true target category. For the extracted features of each target, they can be fed into a classifier (e.g., a softmax classifier) to calculate its corresponding target category, which is the predicted target category.
[0210] In the first round of training, the target's own features are extracted from each training image, and the first target classification loss can be calculated. In subsequent rounds of training, only the first training image in each training sample pair has the target's own features extracted (the target's own features in the second training image are read from the cache, not extracted), and the first target classification loss can be calculated.
[0211] (3) The second target classification loss is calculated based on the contextual features of the target obtained by the feature fusion network.
[0212] The second target classification loss represents the difference between the target category obtained based on the target's contextual features and the true target category (which needs to be labeled in advance). The specific form of the loss function is not limited. This loss can affect the parameter updates in the target detection network, feature extraction network, and feature fusion network.
[0213] In each training round, contextual features of the target are extracted from each training image, allowing for the computation of a second target classification loss. Note that if the contextual features include multiple types of features, such as intra-graph contextual features and cross-graph contextual features, then each type of feature can be used to compute a second target classification loss independently. For example, intra-graph contextual features can be used to compute a corresponding second target classification loss, and cross-graph contextual features can also be used to compute a corresponding second target classification loss.
[0214] In another implementation, in addition to the three losses mentioned above, another loss can be calculated:
[0215] (4) Target identity loss based on the target image processing of the first and second training images in the same training sample pair.
[0216] The target identity loss characterizes the difference between the target identity obtained based on target similarity and the true target identity (which needs to be pre-labeled). Target identity refers to whether the target in the first training image and the target in the second training image are the same target. The specific form of the loss function is not limited, and this loss can affect the parameter updates in the target detection network, feature extraction network, and feature fusion network.
[0217] The similarity between the targets in each training sample pair can only be calculated in non-first round training, and thus the target identity loss can be calculated. However, the target identity loss is not always necessary to calculate. For example, it cannot be calculated if the true target identity is not labeled in the training sample pair. However, the previous losses (1) to (3) are sufficient to train the similarity calculation model.
[0218] For example, in a training sample pair, if target X in the first training image and target Y in the second training image are the same target, then the true target identity can be labeled as 1. If the calculated final similarity between targets X and Y is 0.9, then there is a difference between the two, and the target identity loss can be calculated.
[0219] Note that if multiple similarities are calculated based on the contextual features of the target in the first training image and the contextual features of the target in the second training image, such as contextual similarity, comprehensive similarity, and final similarity, then each similarity can be used to calculate the target identity loss independently. For example, contextual similarity can be used to calculate the corresponding target identity loss, comprehensive similarity can be used to calculate the corresponding target identity loss, and final similarity can also be used to calculate the corresponding target identity loss.
[0220] The above implementation method, by flexibly setting various types of losses, achieves effective supervision of the similarity calculation model, which is beneficial to improving the model's performance. It should be understood that the similarity calculation model can also calculate other losses, or adopt unsupervised training methods; examples will not be provided here. For cases where multiple losses are calculated, these losses can be weighted first to calculate the total loss, and then the parameters of the similarity calculation model can be updated based on the total loss.
[0221] Figure 5 The functional modules included in the image processing apparatus 200 provided in an embodiment of this application are shown. (Refer to...) Figure 5 The image processing apparatus 200 includes:
[0222] The feature acquisition component 210 is used to acquire the self-features of the target in the image to be processed and the self-features of the target in the reference image; wherein, the target in the image to be processed includes the target to be processed, the target in the reference image includes the reference target, and the target to be processed and the reference target are targets whose similarity is to be calculated.
[0223] The feature fusion component 220 is used to fuse the self-features of the target to be processed with the self-features of at least one other target in the image to be processed and / or the reference image to obtain the context features of the target to be processed, and to fuse the self-features of the reference target with the self-features of at least one other target in the image to be processed and / or the reference image to obtain the context features of the reference target.
[0224] The context similarity calculation component 230 is used to calculate the context similarity between the target to be processed and the reference target based on the context features of the target to be processed and the context features of the reference target.
[0225] The final similarity calculation component 240 is used to determine the final similarity between the target to be processed and the reference target based on the contextual similarity between the target to be processed and the reference target.
[0226] The image processing apparatus 200 provided in this application embodiment can be used to implement the methods in the foregoing method embodiments (including any possible implementation of the methods). The implementation principle and the resulting technical effects have been described in the foregoing method embodiments. For the sake of brevity, for any parts not mentioned in the apparatus embodiment, please refer to the corresponding content in the method embodiment.
[0227] Figure 6 This illustration shows a possible structure of the electronic device 300 provided in an embodiment of this application. (Refer to...) Figure 6The electronic device 300 includes a processor 310, a memory 320, and a communication interface 330. These components are interconnected and communicate with each other via a communication bus 340 and / or other forms of connection mechanism (not shown).
[0228] The processor 310 includes one or more (only one is shown in the figure), which can be an integrated circuit chip with signal processing capabilities. The processor 310 can be a general-purpose processor, including a Central Processing Unit (CPU), a Microcontroller Unit (MCU), a Network Processor (NP), or other conventional processors; it can also be a special-purpose processor, including a Graphics Processing Unit (GPU), a Neural-network Processing Unit (NPU), a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Furthermore, when there are multiple processors 310, some can be general-purpose processors and others can be special-purpose processors.
[0229] The memory 320 includes one or more (only one is shown in the figure), which may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0230] Processor 310 and other possible components may access memory 320, reading and / or writing data therein. In particular, one or more computer program instructions may be stored in memory 320, which processor 310 may read and execute to perform the image processing methods and / or image processing techniques provided in the embodiments of this application.
[0231] Communication interface 330 includes one or more (only one is shown in the figure) that can be used to communicate directly or indirectly with other devices to exchange data. Communication interface 330 may include interfaces for wired and / or wireless communication.
[0232] Understandable. Figure 6 The structure shown is for illustrative purposes only; the electronic device 300 may also include components that are more advanced than those shown. Figure 6 The more or fewer components shown, or having the same Figure 6 Different configurations are shown. For example, electronic device 300 may also include an image acquisition module (e.g., a camera) for taking photos or videos, and frames from the captured photos or videos can be used as images to be processed in step S110.
[0233] Figure 6 The components shown can be implemented using hardware, software, or a combination thereof. Electronic device 300 may be a physical device, such as a camera, mobile phone, PC, laptop, tablet, server, or robot, or a virtual device, such as a virtual machine or container. Furthermore, electronic device 300 is not limited to a single device; it can be a combination of multiple devices or a cluster of numerous devices.
[0234] This application also provides a computer-readable storage medium storing computer program instructions. These computer program instructions are read and executed by a processor to perform the image processing method provided in this application. For example, the computer-readable storage medium can be implemented as follows: Figure 6 The memory 320 in the electronic device 300.
[0235] This application also provides a computer program product, which includes computer program instructions. These computer program instructions are read and executed by a processor to perform the image processing method provided in this application.
[0236] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. An image processing method, characterized in that, include: The process involves obtaining the intrinsic features of a target in the image to be processed and the intrinsic features of a target in the reference image; wherein the target in the image to be processed includes the target to be processed, the target in the reference image includes the reference target, and the target to be processed and the reference target are targets whose similarity is to be calculated. The context features of the target to be processed are obtained by fusing the self-features of the target to be processed with the self-features of at least one other target in the image to be processed and the self-features of at least one other target in the reference image; and the context features of the reference target are obtained by fusing the self-features of the reference target with the self-features of at least one other target in the image to be processed and the self-features of at least one other target in the reference image. Calculate the context similarity between the target to be processed and the reference target based on the context features of the target to be processed and the reference target. The final similarity between the target to be processed and the reference target is determined based on the contextual similarity between the target to be processed and the reference target.
2. The image processing method according to claim 1, characterized in that, The contextual feature representation of the target to be processed: the features of the target to be processed obtained when considering the features of the environment surrounding the target.
3. The image processing method according to claim 1 or 2, characterized in that, The step of fusing the self-features of the target to be processed with the self-features of at least one other target in the image to be processed and the self-features of at least one other target in the reference image to obtain the contextual features of the target to be processed includes: The contextual features of the target to be processed are obtained by feature fusion in the following manner: The intra-image context features of the target to be processed are fused with the intra-image context features of at least one other target in the image to be processed. The intra-image context features of the target to be processed are fused with the intra-image context features of at least one target in the reference image to obtain the cross-image context features of the target to be processed. The features of the target to be processed are fused with the features of at least one other target in the image to be processed to obtain the intra-image context features of the target to be processed. The intra-image context features of the target to be processed are fused with the features of at least one target in the reference image to obtain the cross-image context features of the target to be processed. The cross-image context features of the target to be processed are enhanced to obtain the enhanced features of the target to be processed. Among them, the cross-graph context features and enhanced features of the target to be processed are both part of the context features of the target to be processed.
4. The image processing method according to claim 3, characterized in that, The step of calculating the context similarity between the target to be processed and the reference target based on the context features of the target to be processed and the context features of the reference target includes: Each feature similarity is calculated between the feature of the target to be processed and the feature of the same type of the reference target for each type of feature, to obtain at least one feature similarity; wherein each feature similarity corresponds to a feature of one type. The contextual similarity between the target to be processed and the reference target is determined based on the at least one feature similarity.
5. The image processing method according to any one of claims 1-4, characterized in that, Determining the final similarity between the target to be processed and the reference target based on the contextual similarity between the target to be processed and the reference target includes: The similarity between the self-features of the target to be processed and the self-features of the reference target is calculated based on the self-features of the target to be processed and the reference target. Calculate the overall similarity between the target to be processed and the reference target based on the contextual similarity and self-feature similarity between the target to be processed and the reference target; The final similarity between the target to be processed and the reference target is determined based on the comprehensive similarity between the target to be processed and the reference target.
6. The image processing method according to claim 5, characterized in that, Determining the final similarity between the target to be processed and the reference target based on the comprehensive similarity between the target to be processed and the reference target includes: The comprehensive similarity between the target to be processed and the reference target is obtained, as well as the comprehensive similarity between the target to be processed and other targets in the reference image. If the overall similarity between the target to be processed and the reference target is the maximum value among the overall similarities between the target to be processed and all targets in the reference image, then the overall similarity between the target to be processed and the reference target is increased to obtain the final similarity between the target to be processed and the reference target. Alternatively, the overall similarity between the target to be processed and the reference target can be directly determined as the final similarity between the target to be processed and the reference target. If the overall similarity between the target to be processed and the reference target is not the maximum value among the overall similarities between the target to be processed and all targets in the reference image, then the overall similarity between the target to be processed and the reference target is reduced to obtain the final similarity between the target to be processed and the reference target.
7. The image processing method according to claim 6, characterized in that, The step of reducing the overall similarity between the target to be processed and the reference target to obtain the final similarity between the target to be processed and the reference target includes: Based on the comprehensive similarity between the target to be processed and all targets in the reference image, a reduction coefficient corresponding to the reference target is calculated; wherein, the reduction coefficient is less than 1 and is positively correlated with the comprehensive similarity between the target to be processed and the reference target; The final similarity between the target to be processed and the reference target is obtained by multiplying the comprehensive similarity between the target to be processed and the reference target by the reduction coefficient corresponding to the reference target.
8. The image processing method according to any one of claims 1-7, characterized in that, The step of fusing the self-features of the target to be processed with the self-features of at least one other target in the image to be processed and the self-features of at least one other target in the reference image to obtain the contextual features of the target to be processed includes: The feature fusion network is used to fuse the features of the target to be processed with the features of at least one other target in the image to be processed and the features of at least one other target in the reference image to obtain the context features of the target to be processed. The feature fusion network is a neural network, and the structure of the feature fusion network includes one of the following two types: Intra-graph feature fusion units and cross-graph feature fusion units are connected sequentially; The graph-intra-graph feature fusion unit, the cross-graph feature fusion unit, and the feature enhancement unit are connected in sequence.
9. The image processing method according to claim 8, characterized in that, The intra-image feature fusion unit is used to perform weighted fusion of the self-features of the target to be processed with the self-features of at least one other target in the image to be processed, to obtain the intra-image context features of the target to be processed; wherein, the weights used by the self-features in the image to be processed in the intra-image feature fusion unit during weighted fusion are calculated by the intra-image feature fusion unit based on an attention mechanism according to all features undergoing weighted fusion. The cross-graph feature fusion unit is used to perform weighted fusion of the intra-graph context features of the target to be processed with the self-features of at least one target in the reference image to obtain the cross-graph context features of the target to be processed; wherein, the weights used by the self-features in the reference image when performing weighted fusion in the cross-graph feature fusion unit are calculated by the cross-graph feature fusion unit based on an attention mechanism according to all features to be weighted and fused.
10. The image processing method according to any one of claims 1-7, characterized in that, The step of fusing the self-features of the target to be processed with the self-features of at least one other target in the image to be processed and the self-features of at least one other target in the reference image to obtain the contextual features of the target to be processed includes: The feature fusion network is used to fuse the features of the target to be processed with the features of at least one other target in the image to be processed and the features of at least one other target in the reference image to obtain the context features of the target to be processed. The feature fusion network is a neural network, comprising at least one feature fusion module connected in sequence. The first feature fusion module performs feature fusion based on the target's own features in the image to be processed and the target's own features in the reference image, resulting in its own output features of the target in the image to be processed and the target in the reference image. Each feature fusion module other than the first performs feature fusion based on the target's features in the image to be processed and the target's features in the reference image, both output by the previous at least one feature fusion module, resulting in its own output features of the target in the image to be processed and the target in the reference image. The structure of the feature fusion module includes one of the following two options: Intra-graph feature fusion units and cross-graph feature fusion units are connected sequentially; The graph-intra-graph feature fusion unit, the cross-graph feature fusion unit, and the feature enhancement unit are connected in sequence.
11. The image processing method according to any one of claims 8-10, characterized in that, The intrinsic features of the target in the image to be processed and the intrinsic features of the target in the reference image are both calculated using a target detection network and a feature extraction network. Both the target detection network and the feature extraction network are neural networks. The method further includes: Train the target detection network, the feature extraction network, and the feature fusion network; In the first round of training, the target detection network and the feature extraction network are used to calculate the self-features of the targets in all training images, and the calculated self-features are written into the cache. In each subsequent training round, for each training sample pair consisting of two training images, the target detection network and the feature extraction network are used to calculate the target's own features in the first training image. The newly calculated target's own features in the first training image are used to update the target's own features in the first training image already written in the cache, and the target's own features in the second training image are directly read from the cache. In the training sample pair, the first training image corresponds to the image to be processed, and the second training image corresponds to the reference image. Both the first training image and the second training image contain multiple targets, and the second training image in each training sample pair serves as the first training image in at least one other training sample pair.
12. The image processing method according to claim 11, characterized in that, The loss that needs to be calculated during training includes the following three items: The target detection loss is calculated based on the target detection results of the target detection network; wherein, the target detection loss represents the difference between the target detection results and the true target detection results; The first target classification loss is calculated based on the target's own features extracted by the feature extraction network; wherein, the first target classification loss represents the difference between the target category obtained based on the target's own features and the true target category; The second target classification loss is calculated based on the contextual features of the target obtained by the feature fusion network; wherein, the second target classification loss represents the difference between the target category obtained based on the contextual features of the target and the true target category; Or, including the above three items plus: The target identity loss is calculated based on the similarity of targets in the first training image and the second training image in the same training sample pair; the target identity loss characterizes the difference between the target identity obtained based on the similarity of targets and the true target identity, where target identity refers to whether the target in the first training image and the target in the second training image are the same target.
13. The image processing method according to any one of claims 1-12, characterized in that, The method further includes: The image processing task is performed using the final similarity between the target to be processed and the reference target.
14. The image processing method according to claim 13, characterized in that, The image processing task is a target search task, the image to be processed is the image to be searched, the target to be processed is the target to be searched, the reference image is the base image, and the reference target is the labeled target; The step of performing an image processing task using the final similarity between the target to be processed and the reference target includes: The search results are sorted based on the final similarity between the search target and the labeled target, calculated from the image to be searched and each base image, and the search results are determined based on the sorting results.
15. A computer program product, characterized in that, It includes computer program instructions, which, when read and executed by a processor, perform the method as described in any one of claims 1-14.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when read and executed by a processor, perform the method as described in any one of claims 1-14.
17. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores computer program instructions, which are read and executed by the processor to perform the method of any one of claims 1-14.