Image Semantic Matching Method, Apparatus, Device, and Storage Medium
By using depth information to assign weights to the feature map in image semantic matching, a depth feature representation map is generated, and the image similarity and fusion uncertainty perceived information entropy comparison loss value is calculated, the problem of low matching accuracy in the prior art is solved, and more accurate semantic matching is achieved.
Patent Information
- Application Number
- CN202510324729.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-19
AI Technical Summary
Existing image semantic matching methods are difficult to correctly understand the object geometry and make full use of a variety of feature maps, resulting in low matching accuracy.
By obtaining the depth information of the image to be matched, weights are assigned to the feature map, and a depth feature representation map is generated; a target image descriptor is generated based on the depth feature representation map; image similarity and fusion uncertainty perceived information entropy comparison loss value are calculated, and the matching result is determined.
Improve the perception of object geometric structures, find more semantic matching results, and achieve accurate image semantic matching.
Smart Images

Figure CN119850986B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image analysis technology, and particularly to an image semantic matching method, apparatus, device, and storage medium. Background Art
[0002] Existing image semantic matching methods often extract features from images by constructing a feature extractor based on a convolutional neural network. The quality of the features depends on the designed network architecture, and there are problems such as difficulty in correctly understanding the geometric structure of objects and making full use of multiple feature maps, resulting in low matching accuracy. Summary of the Invention
[0003] The main purpose of this application is to provide an image semantic matching method, apparatus, device, and storage medium, aiming to solve the technical problem that existing semantic matching methods have difficulties in correctly understanding the geometric structure of objects and making full use of multiple feature maps, resulting in low matching accuracy.
[0004] To achieve the above object, this application proposes an image semantic matching method, and the image semantic matching method includes:
[0005] Obtain the depth information of the image to be matched, and assign weights to the feature maps of the image to be matched according to the depth information to obtain a depth feature representation map;
[0006] Generate a target image descriptor for the image to be matched based on the depth feature representation map;
[0007] Calculate the image similarity of pixel points in the image to be matched and the fusion uncertainty perception information entropy contrast loss value according to the target image descriptor, and determine the matching result based on the image similarity and the fusion uncertainty perception information entropy contrast value.
[0008] Optionally, the step of generating a target image descriptor for the image to be matched based on the depth feature representation map includes:
[0009] Input the depth feature representation map into a target feature fusion network for feature channel fusion to generate an initial image descriptor;
[0010] Calculate the cosine similarity and Gaussian kernel function similarity of the initial image descriptor, and construct a joint similarity matrix and a fusion uncertainty perception information entropy contrast function;
[0011] Optimize the initial fusion weights in the target feature fusion network based on the joint similarity matrix and the fusion uncertainty perception information entropy contrast function to generate a target image descriptor.
[0012] Optionally, the step of optimizing the initial fusion weights in the target feature fusion network based on the joint similarity matrix and the fusion uncertainty-aware information entropy contrast function to generate a target image descriptor includes:
[0013] Determining a first optimization objective through the fusion uncertainty-aware information entropy contrast loss function;
[0014] Predicting the positions of key points according to the joint similarity matrix, and determining an error loss function through the positions to generate a second optimization objective;
[0015] Optimizing the initial fusion weights in the target feature fusion network based on the first optimization objective and the second optimization objective to generate a target image descriptor.
[0016] Optionally, the step of optimizing the initial fusion weights in the target feature fusion network based on the first optimization objective and the second optimization objective to generate a target image descriptor includes:
[0017] Determining the gradient of the initial fusion weights in the target feature fusion network according to the first optimization objective and the second optimization objective;
[0018] Optimizing the initial fusion weights through the gradient to obtain target fusion weights, and updating the target feature fusion network based on the target fusion weights;
[0019] Inputting the depth feature representation map into the updated target feature fusion network again to generate a target image descriptor.
[0020] Optionally, the step of inputting the depth feature representation map into a target feature fusion network for feature channel fusion to generate an initial image descriptor includes:
[0021] Instantiating the input channels of a preset feature fusion network according to the number of channels of the depth feature representation map to obtain a target feature fusion network;
[0022] Inputting the depth feature representation map into the target feature fusion network for calculation to obtain a feature map to be fused;
[0023] Fusing the feature maps to be fused with the same number of channels through initial fusion weights to generate an initial image descriptor.
[0024] Optionally, the step of obtaining the depth information of the image to be matched and assigning weights to the feature map of the image to be matched according to the depth information to obtain a depth feature representation map includes:
[0025] Perform relative depth estimation on the obtained image to be matched, and obtain the depth information of the image to be matched;
[0026] Obtain a depth estimation feature map according to the depth information, and apply a normalized exponential function to the depth estimation feature map to establish a depth distribution model for each pixel point in the image to be matched, where the depth distribution model is used to calculate the probability that the pixel point falls within each depth interval;
[0027] Apply a maximum index function to the depth distribution model to determine the target depth interval corresponding to each pixel point when the probability is the maximum;
[0028] Calculate the relative depth weight of the image to be matched based on the target depth interval, and perform pixel-by-pixel multiplication of the relative depth weight and the feature map of the image to be matched to obtain a depth feature representation map.
[0029] Optionally, the step of calculating the image similarity and the fusion uncertainty-aware information entropy contrast loss value of the pixel points in the image to be matched according to the target image descriptor, and determining the matching result based on the image similarity and the fusion uncertainty-aware information entropy contrast value includes:
[0030] Calculate the image similarity and the fusion uncertainty-aware information entropy contrast loss function value of the pixel points in the image to be matched according to the target image descriptor;
[0031] Score the pixel points through the image similarity and the fusion uncertainty-aware information entropy contrast loss function value, and select the point with the highest score as the predicted key point for the current semantic match.
[0032] In addition, to achieve the above object, the present application also proposes an image semantic matching device, where the image semantic matching device includes:
[0033] A feature extraction module, configured to obtain the depth information of the image to be matched, and assign weights to the feature map of the image to be matched according to the depth information to obtain a depth feature representation map;
[0034] A feature fusion module, configured to generate a target image descriptor of the image to be matched based on the depth feature representation map;
[0035] A matching output module, configured to calculate the image similarity and the fusion uncertainty-aware information entropy contrast loss value of the pixel points in the image to be matched according to the target image descriptor, and determine the matching result based on the image similarity and the fusion uncertainty-aware information entropy contrast value.
[0036] In addition, to achieve the above object, the present application further provides an image semantic matching device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the image semantic matching method as described above.
[0037] In addition, to achieve the above object, the present application further provides a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the image semantic matching method as described above.
[0038] In the present application, the depth information of the image to be matched is obtained, and weights are assigned to the feature map of the image to be matched according to the depth information to obtain a depth feature representation map; a target image descriptor of the image to be matched is generated based on the depth feature representation map; the image similarity and the fusion uncertainty perception information entropy contrast loss value of the pixel points in the image to be matched are calculated according to the target image descriptor, and the matching result is determined based on the image similarity and the fusion uncertainty perception information entropy contrast value. By embedding the depth information into the feature map, the perception of the object's geometric structure is improved, so as to find more semantic matching results, and then accurately achieve semantic matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0040] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0041] Figure 1 It is a schematic flowchart of the first embodiment of the image semantic matching method of the present application;
[0042] Figure 2 It is a schematic flowchart of matching based on the target image descriptor in the image semantic matching method of the present application;
[0043] Figure 3 It is a schematic flowchart of the second embodiment of the image semantic matching method of the present application;
[0044] Figure 4 It is a schematic diagram of the depth distribution modeling principle of the image semantic matching method of the present application;
[0045] Figure 5Schematic diagram for generating the initial image descriptor of the image semantic matching method of the present application;
[0046] Figure 6 Flowchart of the third embodiment of the image semantic matching method of the present application;
[0047] Figure 7 Flowchart of feature fusion in the target feature fusion network of the present application;
[0048] Figure 8 Schematic diagram of the module structure of the image semantic matching device according to the embodiment of the present application;
[0049] Figure 9 Schematic diagram of the device structure of the hardware operating environment involved in the image semantic matching method according to the embodiment of the present application.
[0050] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0051] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0052] For a better understanding of the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and specific implementation manners.
[0053] The main solution of the embodiment of the present application is: obtaining the depth information of the image to be matched, and assigning weights to the feature map of the image to be matched according to the depth information to obtain a depth feature representation map; generating a target image descriptor of the image to be matched based on the depth feature representation map; calculating the image similarity of the pixel points in the image to be matched and the fusion uncertainty perception information entropy contrast loss value according to the target image descriptor, and determining the matching result based on the image similarity and the fusion uncertainty perception information entropy contrast value.
[0054] Due to the rapid development of autonomous driving technology, precise perception of the vehicle's surrounding environment has become crucial for ensuring driving safety and optimizing driving decisions. Traditional two-dimensional image processing is vulnerable to factors such as illumination changes, occlusion, and dynamic target interference in complex road environments, such as multi-lane, high-speed driving, night-time, or adverse weather conditions, leading to a decline in the accuracy of recognition and matching. Therefore, a pixel-by-pixel matching technology based on semantic correspondence is at the core, enabling the computer to still establish pixel-by-pixel matching results based on semantic information even when the pose and appearance of objects in the image change. However, existing methods often extract features from images by constructing a convolutional neural network-based feature extractor, and the quality of the features depends on the designed network architecture. At the same time, existing feature extractors are difficult to correctly understand the geometric structure of objects in the case of depth ambiguity, resulting in low matching accuracy.
[0055] Therefore, the present application provides a semantic matching method that can perceive the geometric structure of objects and make full use of multiple features. By embedding depth information into the feature map, the perception of the geometric structure of objects is enhanced, thereby finding more semantic matching results and achieving accurate image semantic matching.
[0056] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of implementing the above functions. Hereinafter, the semantic matching unit will be taken as an example to illustrate this embodiment and the following embodiments.
[0057] Based on this, an embodiment of the present application provides an image semantic matching method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the image semantic matching method of the present application.
[0058] In this embodiment, the image semantic matching method includes:
[0059] Step S10, obtaining the depth information of the image to be matched, and assigning weights to the feature map of the image to be matched according to the depth information to obtain a depth feature representation map.
[0060] It should be noted that the images to be matched are two or more pictures that need to be matched in an image semantic matching task. They can be images of the same scene or similar scenes taken from different perspectives, at different times, or under different conditions. For example, in an autonomous driving scenario, the images to be matched can be road images taken by the vehicle's front camera at consecutive time points, or images of the same area taken by cameras at different angles. The objects (such as vehicles, pedestrians, traffic signs, etc.) in these images may vary in appearance, pose, or position, but have certain semantic associations. Depth information describes the position of an object or pixel point in the image in three-dimensional space relative to the camera, indicating the distance of the object or pixel point from the camera. The depth feature representation map is a new form of image representation obtained by integrating depth information into the feature map. It not only contains the feature information in the original feature map but also reflects the influence of depth information on the features, that is, the features in different depth regions have different weights or importance in the map.
[0061] It can be understood that assigning weights to the feature map of the images to be matched according to the depth information can be achieved by dividing the depth information into several intervals. For example, it can be divided into depth intervals such as close range, middle range, and far range according to the characteristics and requirements of the scene. Then, according to the importance of different intervals or the degree of influence on semantic matching, a weight value is set for each interval. In an autonomous driving scenario, the close range area (such as within a few to more than ten meters in front of the vehicle) is more critical for driving decisions and will be assigned a higher weight because the vehicles, pedestrians, etc. in this area directly affect the driving safety of the vehicle; while the weight of the far range area is relatively low. In the subsequent semantic matching process, the depth feature representation map can provide a more discriminative and geometric structure-aware feature representation for feature fusion and matching. For example, in the object recognition task of autonomous driving, the depth feature representation map can help better identify objects such as vehicles and pedestrians at different depth positions, improve the accuracy of semantic matching, and thus provide a more reliable basis for the vehicle's driving decisions (such as avoidance, acceleration, deceleration, etc.).
[0062] Step S20: Generate the target image descriptor of the images to be matched based on the depth feature representation map.
[0063] It should be noted that the target image descriptor is a form of description of the image features generated after operations such as feature channel fusion. It summarizes the key features of the image after integrating depth information and feature fusion processing, and can be used for subsequent operations such as calculating similarity and contrast learning between different images, so as to determine the semantic matching relationship between the images.
[0064] It should be understood that in the process of generating the target image descriptor, the features from different feature extractors (such as self-supervised feature extractors and diffusion model feature extractors) will be involved in the fusion. Different feature extractors can often mine the feature information of images from different perspectives and levels. By adding and fusing the features they extract after the processing of the above-mentioned deep feature representation map and the operation of the feature channel fusion network, various feature resources can be fully utilized, avoiding the limitations of a single feature extraction method, enriching the semantic information contained in the target image descriptor, and enhancing its ability to represent image semantics.
[0065] Step S30: Calculate the image similarity of the pixel points in the image to be matched and the fusion uncertainty-aware information entropy contrast loss value according to the target image descriptor, and determine the matching result based on the image similarity and the fusion uncertainty-aware information entropy contrast value.
[0066] It should be understood that the matching result is the semantic correspondence between the images to be matched determined after a series of calculations and analyses. In the image semantic matching task, it is ultimately necessary to clarify which pixel points or regions correspond to each other in different images. For example, in the autonomous driving scenario, it is necessary to find the corresponding pixel positions of the same vehicle, the same pedestrian, or the same traffic sign, etc. in the road images taken from different perspectives, so as to assist subsequent decision-making judgments, such as target positioning, trajectory analysis, etc.
[0067] It should be understood that when determining the matching result, not only the image similarity index is relied on, but also the two aspects of image similarity and contrast loss value are comprehensively considered. This can avoid misjudgments that may be caused by relying only on a single index. For example, simply looking at the pixel points with high image similarity may result in incorrect matches (due to some similar but semantically non-corresponding parts in the image), and screening by combining the contrast loss value can further eliminate these unreasonable matches and more accurately determine the truly semantically corresponding pixel points, making the final matching result more reliable.
[0068] It can be understood that when determining the matching result based on the image similarity and the fusion uncertainty-aware information entropy contrast loss value, the SoftArgmax function can be used to convert the calculated similarity matrix into a flow field, and then determine the corresponding predicted key point positions of each pixel point on the target image in the source image. At the same time, the pixel points with lower contrast loss values are screened out by combining the fusion uncertainty-aware information entropy contrast loss function value. Finally, the pixel points with the highest similarity score and the lowest fusion uncertainty-aware information entropy contrast loss function loss value, or the key points determined by these pixel points, are used as the predicted key points for semantic matching, so as to determine the semantic matching result between the images to be matched.
[0069] In one example, referring to Figure 2 , Figure 2 is a schematic flowchart of matching based on a target image descriptor in the image semantic matching method of this application. Starting from the target image descriptor, the cosine and Gaussian kernel joint similarity matrix is then calculated. Subsequently, the process is divided into two parallel scoring directions. On the one hand, it is guided by the information entropy of uncertainty perception, and on the other hand, it is based on the average endpoint error of the optical flow field. These two objectives aim to score the image matching from different angles. Finally, based on the results of the above calculations, the matching degree of each pixel point is scored to determine the semantic matching situation of the image.
[0070] Further, in order to more accurately determine the most likely matching key points in the image, the step S30 may include:
[0071] Calculating the image similarity of the pixel points in the image to be matched and the value of the fusion uncertainty perception information entropy contrast loss function according to the target image descriptor; scoring the pixel points through the image similarity and the value of the fusion uncertainty perception information entropy contrast loss function, and selecting the point with the highest score from the pixel points as the predicted key point of the current semantic matching.
[0072] In one example, the image processing module in the autonomous driving system performs feature extraction and feature fusion operations on the collected road images A and B to generate corresponding target image descriptors. The target image descriptors contain rich feature information, including features such as the shape, color, and position of a car in front in the image, as well as the texture features of the road. Then, the semantic matching unit calculates the image similarity of pixel points in the two images. For each pixel point in the image, taking pixel point a as an example, the cosine similarity and Gaussian kernel function similarity are calculated between its corresponding target image descriptor vector and the target image descriptor vectors of all pixel points in image B. At the same time, the obtained similarity matrix is applied to the fusion uncertainty perception information entropy contrast loss function to calculate the fusion uncertainty perception information entropy contrast loss value. After obtaining the image similarity matrix and the loss value of the fusion uncertainty perception information entropy contrast loss function, the prediction key points of semantic matching start to be screened. All pixel points in images A and B are traversed to find those pixel points with the highest image similarity score and the lowest loss value of the fusion uncertainty perception information entropy contrast loss function. Suppose in images A and B, it is found that for pixel point b representing the center position of the left front wheel of a certain car in front, its similarity with a certain pixel point in image B is the highest among all the pixel points compared with it, reaching 0.95. At the same time, the loss value of the fusion uncertainty perception information entropy contrast loss function corresponding to this pair of pixel points is also extremely low, which is more in line with the optimization requirements compared with other pixel point pairs with high similarity. Then pixel point b is the corresponding position of the center of the left front wheel of this car in the two images, and this pixel point can be determined as the prediction key point of semantic matching.
[0073] In this embodiment, the depth information of the image to be matched is obtained, and weights are assigned to the feature map of the image to be matched according to the depth information to obtain a depth feature representation map; a target image descriptor of the image to be matched is generated based on the depth feature representation map; the image similarity and the fusion uncertainty perception information entropy contrast loss value of pixel points in the image to be matched are calculated according to the target image descriptor, and the matching result is determined based on the image similarity and the fusion uncertainty perception information contrast value. By embedding the depth information into the feature map, the perception of the geometric structure of the object is improved, so as to find more semantic matching results, and then accurate semantic matching is realized.
[0074] Refer to Figure 3 , Figure 3 FIG.
[0075] In the second embodiment, step S10 includes:
[0076] Step S101: Perform relative depth estimation on the obtained image to be matched, and obtain the depth information of the image to be matched.
[0077] It should be noted that relative depth estimation is a process of inferring the relative depth information of pixel points in an image through image analysis techniques, that is, determining the distance relationship of different objects or regions in the image relative to the camera in three-dimensional space. Based on various visual cues in the image, such as texture gradients, light and shadow changes, object sizes, and perspective relationships, etc., to estimate the depth order or relative depth range of each pixel point.
[0078] It can be understood that performing relative depth estimation on the obtained image to be matched can conduct a comprehensive visual cue analysis on the image to be matched. For example, texture gradients can provide information about the changes in the object surface in the depth direction; the distribution of the light-receiving surface and the backlight surface of the object can imply its shape and position relationship, and then infer the depth; according to the perspective principle, distant objects appear smaller in the image, while nearby objects are relatively larger. By comparing the size ratio relationships of different objects in the image, their relative depth order can be initially determined. After analyzing various visual cues, relevant features need to be extracted from the image and these features are integrated to obtain depth information. In this process, the depth information provided by different cues will be weighted and fused to obtain a more accurate and reliable depth estimation result. For example, for areas with clear texture and obvious light and shadow changes, a higher weight is given because these areas can usually provide more accurate depth information.
[0079] Step S102: Obtain a depth estimation feature map according to the depth information, and apply the sigmoid function to the depth estimation feature map to establish a depth distribution model for each pixel point in the image to be matched, where the depth distribution model is used to calculate the probability that the pixel point falls within each depth interval.
[0080] It should be noted that the sigmoid function is to convert a real number vector into a vector representing the probability distribution of each category, so that each element in the vector is between 0 and 1, and the sum of all elements is 1. The depth distribution model describes the probability situation that the depth value of this pixel point falls within each pre-set depth interval. For example, several different depth intervals are pre-set, such as the near-view interval, the middle-view interval, the far-view interval, etc., and the possibility of each pixel point being in these intervals is calculated respectively.
[0081] Step S103: Apply the argmax function to the depth distribution model to determine the target depth interval corresponding to each pixel point when the probability is the largest.
[0082] It should be noted that the maximum index function returns the index value corresponding to the maximum value of a certain objective function in a given set of values. In the application scenario of the depth distribution model, the input is the probability vector of each depth interval corresponding to each pixel point, and the output is the index corresponding to the depth interval with the highest probability. Through this index, the depth interval where the pixel point is most likely to be located can be determined.
[0083] In one example, referring to Figure 4 , Figure 4 is a schematic diagram of the depth distribution modeling principle of the image semantic matching method of this application. For the image of the road in front of the vehicle, after determining the target depth interval corresponding to each pixel point through the index function, different weights can be assigned to the pixel points in different depth intervals (for example, higher weights are assigned to the pixel points in the depth intervals where nearby vehicles, pedestrians, etc. are located), which are used for subsequent image feature processing and semantic matching to assist the autonomous driving system in making reasonable driving decisions, such as braking in time to avoid nearby obstacles. The left part of the figure shows a scene of the road in front of a vehicle, including elements such as vehicles, pedestrians, and traffic lights. In this scene, through a certain index function, the target depth interval corresponding to each pixel point can be determined. The middle part of the image shows the annotation of different targets. For example, vehicles and pedestrians are marked with frames of different colors, indicating that different targets can be recognized when processing the image. The right part of the figure shows the weight distribution of different depth intervals. For example, the red curve represents the weight distribution of nearby targets, and its peak is in a relatively close depth interval, which means that the system assigns a higher weight to nearby targets. The green and blue curves may represent the weight distributions of targets at different distances, and their peaks are in relatively far depth intervals. Through this depth distribution modeling, in the autonomous driving scenario, the system can process image features more accurately. Higher weights are assigned to the pixel points in the depth intervals where nearby vehicles and pedestrians are located, so that the system can pay more attention to these key targets when performing semantic matching and decision-making. For example, when the system detects a pedestrian or vehicle nearby, it can make reasonable driving decisions such as braking in time to avoid, thereby improving the safety and reliability of autonomous driving.
[0084] Step S104, calculate the relative depth weight of the image to be matched based on the target depth interval, and multiply the relative depth weight and the feature map of the image to be matched pixel by pixel to obtain a depth feature representation map.
[0085] It should be noted that the relative depth weight is the weight value assigned to the pixel points in the image to be matched according to the target depth interval. Usually, the pixel points corresponding to the depth interval closer to the camera are assigned higher weights because they have a greater impact on the overall image understanding and related decisions; while the weights of the pixel points farther from the camera are relatively lower. By setting such weights, the importance differences of features in different depth regions are highlighted.
[0086] It can be understood that, according to the image to be matched, a depth estimation feature map can be obtained based on depth estimation, the normalized exponential function is applied to the depth estimation feature map, and a depth distribution model is established for each pixel. The interval of the depth distribution is predefined as D discrete intervals, such as . Wherein the depth distribution describes the probability that the depth value of the pixel falls within each depth interval. Apply function to obtain the most likely depth interval of each pixel in the depth distribution, and obtain the relative depth weight of the image to be matched. Then, the relative depth weight is multiplied pixel by pixel with the feature map of the image to be matched obtained from the self-supervised feature extractor and the diffusion model feature extractor to obtain a depth feature representation map weighted by depth information.
[0087] In one example, refer to Figure 5 , Figure 5 is a schematic diagram for generating an initial image descriptor of the image semantic matching method of the present application. First, the image to be matched is respectively processed by a feature extractor and a relative depth estimator to obtain a feature map and depth distribution modeling. The depth weight is obtained through depth distribution modeling. The feature map is combined into a depth feature representation map based on the depth weight, and the depth feature representation map is further processed by a feature fusion network. During the processing, based on an optimization method guided by uncertainty perception and information entropy, an initial image descriptor is finally generated.
[0088] In this embodiment, relative depth estimation is performed on the obtained image to be matched to obtain the depth information of the image to be matched; a depth estimation feature map is obtained according to the depth information, and the normalized exponential function is applied to the depth estimation feature map to establish a depth distribution model for each pixel point in the image to be matched, wherein the depth distribution model is used to calculate the probability that the pixel point falls within each depth interval; the maximum index function is applied to the depth distribution model to determine the target depth interval corresponding to each pixel point when the probability is the largest; the relative depth weight of the image to be matched is calculated based on the target depth interval, and the relative depth weight is multiplied pixel by pixel with the feature map of the image to be matched to obtain a depth feature representation map. Using the depth estimation feature map to establish a depth distribution model to describe the possibility of each pixel point in different depth intervals in the form of probability can more accurately grasp their depth distribution, and enhance the expression ability of the depth feature representation map for the object geometric structure.
[0089] Refer to Figure 6 , Figure 6 is a schematic flowchart of the third embodiment of the image semantic matching method of the present application. Based on the above second embodiment, the third embodiment of the image semantic matching method of the present application is proposed.
[0090] In the third embodiment, step S20 includes:
[0091] Step S201: Input the depth feature representation map into a target feature fusion network for feature channel fusion to generate an initial image descriptor.
[0092] It should be noted that the target feature fusion network is a network structure used for feature fusion operations on the depth feature representation map, usually composed of multiple layers with different functions. The initial image descriptor is a form of description of image features generated after operations such as feature channel fusion, which summarizes the key features of the image after integrating depth information and feature fusion processing, and can be used for subsequent operations such as calculating similarity and contrast learning between different images, so as to determine the semantic matching relationship between images.
[0093] It can be understood that the target feature fusion network needs to utilize the characteristics of the convolutional layer to extract and fuse features during the feature fusion process. The depth convolutional layer performs feature fusion in the depth direction through a relatively large convolutional kernel (such as size 7) to capture more macroscopic feature relationships; while the 1×1 convolutional layer is used to flexibly adjust the number of channels and refine the feature representation, and cooperate with the activation layer to introduce non-linearity, enabling the network to learn complex and diverse image semantic features, and realizing effective feature fusion through such progressive convolutional operations.
[0094] Further, in order to achieve the adaptive matching between the network structure and the input data, step S201 may include:
[0095] Instantiate the input channels of a preset feature fusion network according to the number of channels of the depth feature representation map to obtain a target feature fusion network; input the depth feature representation map into the target feature fusion network for calculation to obtain a to-be-fused feature map; fuse the to-be-fused feature maps with the same number of channels through an initial fusion weight to generate an initial image descriptor.
[0096] It should be noted that the number of channels is the quantity used to describe image features from different dimensions or perspectives. The initial fusion weight is a weight parameter initially set during the feature fusion process to determine the proportion of different to-be-fused feature maps during fusion.
[0097] It should be understood that the initial fusion weight can be randomly initialized, and then continuously adjusted through the backpropagation algorithm according to specific optimization objectives (such as the fusion uncertainty-aware information entropy contrast loss function, etc.) during the subsequent training process, so that the fused image descriptor can better reflect the semantic information of the image and improve the accuracy of semantic matching.
[0098] In one example, refer to Figure 7 , Figure 7This is the flowchart of feature fusion in the target feature fusion network of this application. First, according to the number of channels of the input k feature representation graphs, k feature fusion networks with different numbers of input channels are instantiated. The number of output channels of the feature fusion network can be unified to 768 to facilitate subsequent descriptor operations on the images to be matched. The feature fusion network consists of, from the first layer to the last layer: a depth convolution layer with a convolution kernel size of 7, a convolution layer with a convolution kernel size of 1, an activation layer, and a convolution layer with a convolution kernel size of 1. After obtaining the feature maps to be fused with the same number of channels, multiply the feature maps to be fused by the learnable fusion weights of this feature fusion network , and then add the k features with aligned channel numbers to obtain the target image descriptor of the image to be matched.
[0099] Step S202: Calculate the cosine similarity and Gaussian kernel function similarity of the initial image descriptor, and construct a joint similarity matrix and a fusion uncertainty-aware information entropy contrast function.
[0100] It should be noted that in image semantic matching, image features are usually represented in vector form. The cosine similarity is used to measure the similarity between two image feature vectors. The Gaussian kernel function similarity is used to measure the similarity of image features in space. The joint similarity matrix is a matrix obtained by combining the cosine similarity and the Gaussian kernel function similarity according to certain rules. Each element in the matrix corresponds to the joint similarity between two image features. The fusion uncertainty-aware information entropy contrast function is used to evaluate and optimize the matching result by comparing the information entropy during the image matching process to cope with the uncertainty and noise in the image data.
[0101] In one example, calculate the cosine similarity and Gaussian kernel function of the image descriptor to be matched. By weighted fusion of the cosine similarity and the kernel function-based similarity, the characteristics of both can be utilized simultaneously to construct a new joint similarity metric. Specifically:
[0102]
[0103] Among them, represents the cosine similarity matrix, is the image descriptor to be matched, represents the source image, represents the target image, represents different pixel points of the descriptor.
[0104]
[0105] Among them, represents the Gaussian kernel function similarity matrix, is the smoothing parameter of the Gaussian kernel.
[0106] Then the two similarity matrices are weighted and fused into a joint similarity matrix :
[0107]
[0108] At this point, uncertainty is introduced to measure the credibility of the model prediction. A noise parameter is added based on the joint similarity matrix, assuming that the noise conforms to the normal distribution:
[0109]
[0110] in, represents the variance corresponding to the uncertainty.
[0111] Probability distribution considering uncertainty for:
[0112]
[0113]
[0114] in, Indicates the source image The strength of a potential match between a location and the target image. Indicates the target image The strength of a potential match between a location and the source image.
[0115] Calculate the information entropy contrast loss function :
[0116]
[0117]
[0118] Based on the similarity calculation of the initial image descriptor and the application of the fusion uncertainty-aware information entropy comparison function, the uncertainty can be optimized by optimizing the probability distribution of uncertainty and feature similarity. Contrastive learning can shorten the feature distance of similar samples and extend the feature distance of dissimilar samples, thereby optimizing the feature embedding space and making the subsequent matching process more accurate and stable.
[0119] Step S203, optimizing the initial fusion weights in the target feature fusion network based on the joint similarity matrix and the fusion uncertainty-aware information entropy comparison function to generate a target image descriptor.
[0120] It should be understood that the initial fusion weights in the target feature fusion network can be optimized by training all images and key points in a preset data set, and when the training is completed until the network converges, the network weights after the training are saved.
[0121] Furthermore, in order to guide the network to learn a feature embedding space in the training process that makes the features of similar samples closer and the features of dissimilar samples farther away, and enhance the accuracy and stability of semantic matching. The step S203 may include:
[0122] Determine a first optimization target through the fusion uncertainty-aware information entropy contrast loss function; predict the position of the key points according to the joint similarity matrix, and generate a second optimization target through the position determination error loss function; optimize the initial fusion weights in the target feature fusion network based on the first optimization target and the second optimization target to generate a target image descriptor.
[0123] It can be understood that predicting the position of the key points according to the similarity matrix can optimize the flow consistency based on the consistency loss of optical flow matching, and use the Euclidean distance between the predicted key points and the true key points as the average endpoint error loss function.
[0124] In one example, calculate the contrast loss from the source image to the target image and from the target image to the source image respectively, and add the two as the first optimization target: The calculation formula is as follows:
[0125] =
[0126]
[0127]
[0128] Among them, 1 represents the contrast loss from the source image to the target image, 2 represents the contrast loss from the target image to the source image, and N represents the total number of samples. is the weight, which determines the participation degree of the information entropy. The derivation formula of 1 is:
[0129] 1=
[0130] Among them, represents different pixel points of the descriptor, represents the positive sample pixel points with the same label as and N represents the total number of samples.
[0131] Correspondingly, The derivation formula of 2 is:
[0132] 2=
[0133] For semantic matching, in order to obtain more accurate matching results, it is necessary to minimize the value of the contrast loss.
[0134] Take the Euclidean distance between the key points and the true key points as the average endpoint error loss function to determine the second optimization objective: , based on the gradient of the obtained second optimization objective, the initial fusion weight parameters in the target feature fusion network can be optimized using the backpropagation principle. Among them The calculation process is as follows:
[0135]
[0136]
[0137] Among them, is the joint similarity matrix with added uncertainty, is the corresponding predicted key point on the target image, represents the true key point, is the regularization term of uncertainty, which controls the size of uncertainty during prediction.
[0138] Further, in order to enable the updated network to better capture the semantic features of objects in the image and improve the adaptability to different scenarios and targets. The step of optimizing the initial fusion weights in the target feature fusion network based on the first optimization objective and the second optimization objective to generate a target image descriptor may include:
[0139] Determine the gradient of the initial fusion weights in the target feature fusion network according to the first optimization objective and the second optimization objective; optimize the initial fusion weights through the gradient to obtain the target fusion weights, and update the target feature fusion network based on the target fusion weights; input the depth feature representation map into the updated target feature fusion network again to generate a target image descriptor.
[0140] In an example, first, by weighting the two optimization objective functions, the total optimization function is determined:
[0141] =ɑ +β
[0142] Among them, ɑ and β are pre-set optimization weight coefficients.
[0143] Then calculate the gradient of the total optimization function with respect to the initial fusion weights in the target feature fusion network. If the initial fusion weights are represented by a vector , where n is the number of weights, then the gradient of the initial fusion weights is:
[0144]
[0145] According to the calculated gradient, use an optimization algorithm to update the initial fusion weights. Use AdamW as the optimizer and a cosine annealing learning rate scheduler. Set the learning rate to 1e-3 and the batch size to 2 to calculate the updated target fusion weights. Apply the obtained target fusion weights to the target feature fusion network to replace the original initial fusion weights and complete the update of the network. Then, input the previously obtained depth feature representation map into the updated target feature fusion network again to generate the target image descriptor.
[0146] In this embodiment, input the depth feature representation map into the target feature fusion network for feature channel fusion to generate an initial image descriptor; calculate the cosine similarity and Gaussian kernel function similarity of the initial image descriptor, and construct a joint similarity matrix and a fusion uncertainty-aware information entropy contrast function; optimize the initial fusion weights in the target feature fusion network based on the joint similarity matrix and the fusion uncertainty-aware information entropy contrast function to generate the target image descriptor. Determine the fusion uncertainty-aware information entropy contrast loss function and the similarity matrix based on the joint similarity matrix, and further determine the optimization target through the fusion uncertainty-aware information entropy contrast loss function to guide the network to learn a feature embedding space in the training process that makes similar sample features closer and dissimilar sample features farther away, enhancing the accuracy and stability of semantic matching.
[0147] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the image semantic matching method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0148] This application also provides an image semantic matching device. Please refer to Figure 8 , the image semantic matching device includes:
[0149] A feature extraction module 10, configured to obtain the depth information of the image to be matched, and assign weights to the feature map of the image to be matched according to the depth information to obtain a depth feature representation map;
[0150] A feature fusion module 20, configured to generate a target image descriptor of the image to be matched based on the depth feature representation map;
[0151] A matching output module 30 is configured to calculate the image similarity of pixel points in the image to be matched and the fusion uncertainty perception information entropy contrast loss value according to the target image descriptor, and determine a matching result based on the image similarity and the fusion uncertainty perception information entropy contrast value.
[0152] The image semantic matching device provided in this application adopts the image semantic matching method in the above embodiment, and can solve the technical problem that the existing semantic matching method has difficulty in correctly understanding the geometric structure of an object and making full use of multiple feature maps, resulting in a low matching accuracy. Compared with the prior art, the beneficial effects of the image semantic matching device provided in this application are the same as those of the image semantic matching method provided in the above embodiment, and other technical features in the image semantic matching device are the same as the features disclosed in the method of the above embodiment, which will not be elaborated here.
[0153] This application provides an image semantic matching device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the image semantic matching method in the first embodiment above.
[0154] Next, refer to Figure 9 , which shows a schematic structural diagram of an image semantic matching device suitable for implementing the embodiments of this application. The image semantic matching device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistant), PADs (Portable Application Description), PMPs (Portable Media Player), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The image semantic matching device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of this application.
[0155] As Figure 9As shown in the figure, the image semantic matching device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the image semantic matching device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the image semantic matching device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an image semantic matching device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.
[0156] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0157] The image semantic matching device provided by the present application adopts the image semantic matching method in the above-mentioned embodiment, and can solve the technical problems that the existing semantic matching methods have difficulty in correctly understanding the geometric structure of objects and making full use of multiple feature maps, resulting in low matching accuracy. Compared with the prior art, the beneficial effects of the image semantic matching device provided by the present application are the same as those of the image semantic matching method provided by the above-mentioned embodiment, and other technical features in the image semantic matching device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0158] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0159] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0160] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the image semantic matching method in the above embodiments.
[0161] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM: Random Access Memory), read-only memory (ROM: Read Only Memory), erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0162] The above computer-readable storage medium can be included in the image semantic matching device; or it can exist separately without being assembled into the image semantic matching device.
[0163] The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by the image semantic matching device, the image semantic matching device is caused to execute the image semantic matching method described above.
[0164] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN: Local Area Network) or a wide area network (WAN: Wide Area Network), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0165] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutively represented blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0166] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0167] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned image semantic matching method, which can solve the technical problem that the existing semantic matching methods are difficult to correctly understand the geometric structure of objects and make full use of multiple feature maps, resulting in low matching accuracy. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the image semantic matching method provided by the above embodiments, and will not be elaborated here.
[0168] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields is included in the patent protection scope of the present application.
Claims
1. An image semantic matching method, characterized in that: The image semantic matching method comprises: Acquire depth information of the image to be matched, and assign weights to the feature map of the image to be matched according to the depth information to obtain a depth feature representation map; Generate a target image descriptor of the image to be matched based on the deep feature representation graph; The image similarity and the fused uncertainty-perceived information entropy contrast loss value of the pixel points in the image to be matched are calculated according to the target image descriptor, and the matching result is determined based on the image similarity and the fused uncertainty-perceived information entropy contrast loss value, wherein the fused uncertainty-perceived information entropy contrast loss value fuses the cosine similarity and the Gaussian kernel function similarity, and calculates the potential matching strength between each position in the source image and the target image by introducing uncertainty.
2. The image semantic matching method according to claim 1, characterized in that: The step of generating a target image descriptor of the image to be matched based on the deep feature representation graph comprises: Inputting the deep feature representation graph into a target feature fusion network to perform feature channel fusion to generate an initial image descriptor; Calculating the cosine similarity and Gaussian kernel function similarity of the initial image descriptor, constructing a joint similarity matrix and a fusion uncertainty perception information entropy comparison loss function; The initial fusion weights in the target feature fusion network are optimized based on the joint similarity matrix and the fusion uncertainty-aware information entropy contrast loss function to generate a target image descriptor.
3. The image semantic matching method according to claim 2, characterized in that: The step of optimizing the initial fusion weights in the target feature fusion network based on the joint similarity matrix and the fusion uncertainty perception information entropy contrast loss function to generate a target image descriptor includes: Determining a first optimization objective by comparing the fused uncertainty-aware information entropy loss function; Predicting the positions of key points according to the joint similarity matrix, and determining an error loss function through the positions to generate a second optimization objective; The initial fusion weights in the target feature fusion network are optimized based on the first optimization objective and the second optimization objective to generate a target image descriptor.
4. The image semantic matching method according to claim 3, characterized in that: The step of optimizing the initial fusion weights in the target feature fusion network based on the first optimization objective and the second optimization objective to generate a target image descriptor comprises: Determining the gradient of the initial fusion weight in the target feature fusion network according to the first optimization objective and the second optimization objective; Optimizing the initial fusion weights by using the gradients to obtain target fusion weights, and updating the target feature fusion network based on the target fusion weights; The deep feature representation map is input again into the updated target feature fusion network to generate a target image descriptor.
5. The image semantic matching method according to claim 2, characterized in that: The step of inputting the deep feature representation map into a target feature fusion network to perform feature channel fusion to generate an initial image descriptor comprises: Instantiate the entry channel of the preset feature fusion network according to the number of channels of the deep feature representation map to obtain a target feature fusion network; Inputting the deep feature representation graph into the target feature fusion network for calculation to obtain a feature graph to be fused; The feature maps to be fused with the same number of channels are fused through initial fusion weights to generate an initial image descriptor.
6. The image semantic matching method according to claim 1, characterized in that: The step of obtaining the depth information of the image to be matched, and assigning a weight to the feature map of the image to be matched according to the depth information to obtain a depth feature representation map includes: Performing relative depth estimation on the obtained image to be matched to obtain depth information of the image to be matched; Obtaining a depth estimation feature map according to the depth information, and applying a normalized exponential function to the depth estimation feature map to establish a depth distribution model for each pixel in the image to be matched, wherein the depth distribution model is used to calculate the probability that the pixel falls within each depth interval; Applying a maximum index function to the depth distribution model to determine a target depth interval corresponding to each of the pixel points when the probability is maximum; The relative depth weight of the image to be matched is calculated based on the target depth interval, and the relative depth weight is multiplied pixel by pixel by the feature map of the image to be matched to obtain a depth feature representation map.
7. The image semantic matching method according to any one of claims 1 to 6, characterized in that: The step of calculating the image similarity and the fused uncertainty perception information entropy contrast loss value of the pixel points in the to-be-matched image according to the target image descriptor, and determining the matching result based on the image similarity and the fused uncertainty perception information entropy contrast loss value comprises: Calculate the image similarity and fusion uncertainty perception information entropy contrast loss value of the pixel points in the to-be-matched image according to the target image descriptor; The pixel points are scored by using the image similarity and the fused uncertainty-aware information entropy contrast loss value, and the point with the highest score is selected from the pixel points as the predicted key point of the current semantic matching.
8. An image semantic matching device, characterized in that: The device comprises: A feature extraction module, used to obtain depth information of the image to be matched, and assign weights to the feature map of the image to be matched according to the depth information to obtain a depth feature representation map; A feature fusion module, used to generate a target image descriptor of the image to be matched based on the deep feature representation graph; A matching output module is used to calculate the image similarity and the fused uncertainty-perceived information entropy contrast loss value of the pixel points in the to-be-matched image according to the target image descriptor, and determine the matching result based on the image similarity and the fused uncertainty-perceived information entropy contrast loss value, wherein the fused uncertainty-perceived information entropy contrast loss value fuses the cosine similarity and the Gaussian kernel function similarity, and calculates the potential matching strength between each position in the source image and the target image by introducing uncertainty.
9. An image semantic matching device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the image semantic matching method according to any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the image semantic matching method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Image matching method and device, equipment and storage medium
CN110781911A
Image feature matching method, system and device and computer storage medium
CN119251527A