Image matching method, device, storage medium and electronic device

By extracting the features of the image to be matched and the reference image in image matching, calculating the feature correlation coefficient, and determining the matching area, the problem of low image matching efficiency in the prior art is solved, and an efficient and general image matching method is realized.

CN114548218BActive Publication Date: 2025-05-13NETEASE (SHANGHAI) NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210032407.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-12
Publication Date
2025-05-13
Estimated Expiration
2042-01-12

AI Technical Summary

Technical Problem

In the prior art, the image matching method is inefficient, requires a lot of labor and time costs, and the generalization performance of the model is difficult to predict and it is difficult to produce a general model.

Method used

By obtaining the image to be matched and its corresponding reference image, the features between the two are extracted, and the target correlation coefficient between the features is calculated, and the matching region is determined in the reference image based on the correlation coefficient.

Benefits of technology

This method improves the efficiency of image matching, avoids the large amount of data sets and time costs required based on supervised learning methods, and meets the generality conditions in automated testing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114548218B_ABST
    Figure CN114548218B_ABST
Patent Text Reader

Abstract

The present invention discloses an image matching method, device, storage medium and electronic device. The method comprises: obtaining a first image and a second image, wherein the first image is an image to be matched, and the second image is a reference image corresponding to the image to be matched; extracting a first feature from the first image, and extracting a second feature from the second image; obtaining a first target correlation coefficient between the first feature and the second feature, wherein the first target correlation coefficient is used to indicate the degree of correlation between the first feature and the second feature; and determining a first target area matching the first image in the second image based on the first target correlation coefficient. The present invention solves the technical problem of low efficiency of image matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to an image matching method, device, storage medium and electronic device. Background Art

[0002] At present, when performing image matching, an image matching method based on supervised learning can be used. This method requires the pre-production of a certain scale of data sets, the training of models through the data sets, and the image matching based on the models. However, this method requires a lot of manpower and time costs in the early stage, and the generalization performance of the model is difficult to estimate, and it is difficult to produce a general model, so there is a technical problem of low efficiency of image matching.

[0003] Currently, no effective solution has been proposed to the above-mentioned technical problem of low efficiency of image matching. Summary of the invention

[0004] At least some embodiments of the present invention provide an image matching method, apparatus, storage medium and electronic device to at least solve the technical problem of low efficiency of image matching.

[0005] According to one embodiment of the present invention, an image matching method is provided. The method may include: acquiring a first image and a second image, wherein the first image is an image to be matched, and the second image is a reference image corresponding to the image to be matched; extracting a first feature from the first image, and extracting a second feature from the second image; acquiring a first target correlation coefficient between the first feature and the second feature, wherein the first target correlation coefficient is used to indicate the degree of correlation between the first feature and the second feature; and determining a first target region matching the first image in the second image based on the first target correlation coefficient.

[0006] Optionally, extracting a first feature from the first image includes: extracting a first target feature map from the first image based on a feature extraction model, wherein the feature extraction model is obtained based on convolutional neural network training, and the first target feature map includes: shape features of the first image.

[0007] Optionally, the method further includes: outputting a first target feature map based on a feature output layer of the feature extraction model, wherein the feature output layer is determined based on a size of the first image.

[0008] Optionally, extracting the second feature from the second image includes: extracting a second target feature map from the second image based on a feature extraction model, wherein the second target feature map includes: shape features of the second image.

[0009] Optionally, the method further includes: outputting a second target feature map based on a feature output layer of the feature extraction model, wherein the feature output layer is determined based on a size of the first image.

[0010] Optionally, obtaining a first target correlation coefficient between the first feature and the second feature includes: transforming the first target feature map from the time domain to the frequency domain to obtain a third target feature map; transforming the second target feature map from the time domain to the frequency domain to obtain a fourth target feature map; performing normalized cross-correlation processing on the third target feature map and the fourth target feature map to obtain the first target correlation coefficient.

[0011] Optionally, normalized cross-correlation processing is performed on the third target feature map and the fourth target feature map to obtain a first target correlation coefficient, including: determining a first complex conjugate value corresponding to the third target feature map and a second complex conjugate value corresponding to the fourth target feature map; performing an inverse Fourier transform on the product of the first complex conjugate value and the second complex conjugate value to obtain the first target correlation coefficient.

[0012] Optionally, the first target feature map is a multi-channel first target feature map, the second target feature map is a multi-channel second target feature map, and obtaining the first target correlation coefficient between the first feature and the second feature includes: obtaining the first target correlation coefficient between the first target feature map of each channel and the second target feature map of each channel to obtain multiple first target correlation coefficients; determining a first target area matching the first image in the second image based on the first target correlation coefficient, including: determining the first target area in the second image based on the multiple first target correlation coefficients.

[0013] Optionally, determining the first target area in the second image based on multiple first target correlation coefficients includes: obtaining a maximum first target correlation coefficient among multiple first target correlation coefficients; adjusting the maximum first target correlation coefficient based on a target adjustment parameter; determining a correlation coefficient among the multiple first target correlation coefficients that is greater than or equal to the adjusted maximum first target correlation coefficient as a second target correlation coefficient; and determining the first target area in the second image based on the second target correlation coefficient.

[0014] Optionally, determining the first target area in the second image based on the second target correlation coefficient includes: determining first position information corresponding to the second target correlation coefficient; determining second position information in the second image based on the first position information; and determining the first target area based on the second position information.

[0015] Optionally, determining the second position information in the second image based on the first position information includes: converting the first position information into the second position information based on the width and height of the first target feature map, the width and height of the second target feature map, and the scaling ratio of the second image to the second target feature map.

[0016] Optionally, determining the first target area based on the second position information includes: determining the second position information as the position information of the center of the first target area; and determining a bounding box of the first target area based on the position information of the center to obtain the first target area.

[0017] Optionally, in the case where the number of second target correlation coefficients is multiple, the number of bounding boxes is multiple, and the bounding box of the first target area is determined based on the position information of the center to obtain the first target area, including: selecting a target bounding box from the multiple bounding boxes based on the intersection-and-union ratio between the multiple bounding boxes; determining that the number of target bounding boxes is multiple, then selecting a first target bounding box from the multiple target bounding boxes based on the target points in the second image; and determining the area of ​​the first target bounding box in the second image as the first target area.

[0018] Optionally, the receptive field size of the network layer of the feature extraction model does not exceed the size of the second image.

[0019] Optionally, the method also includes: if it is determined that the first image includes first text information, the first text information is extracted from the first image, and the second text information is extracted from the second image; fuzzy matching is performed on the first text information and the second text information; if it is determined that the fuzzy matching of the first text information and the second text information is successful, a second target area matching the first text information is determined in the second image.

[0020] Optionally, determining the second target area matching the first text information in the second image includes: determining third position information of the first text information in the second image; and determining the second target area based on the third position information.

[0021] Optionally, extracting the first feature from the first image and extracting the second feature from the second image includes: determining that the first image does not include the first text information, or determining that fuzzy matching of the first text information and the second text information fails, then extracting the first feature from the first image and extracting the second feature from the second image.

[0022] Optionally, the method further includes: obtaining a similarity between an image corresponding to the first target area and the first image; and outputting a prompt message if the similarity is greater than a target threshold, wherein the prompt message is used to indicate that the first image matches the second image successfully.

[0023] According to one embodiment of the present invention, an image matching device is also provided. The device may include: a first acquisition unit, used to acquire a first image and a second image, wherein the first image is an image to be matched, and the second image is a reference image corresponding to the image to be matched; an extraction unit, used to extract a first feature from the first image, and extract a second feature from the second image; a second acquisition unit, used to acquire a first target correlation coefficient between the first feature and the second feature, wherein the first target correlation coefficient is used to indicate the degree of correlation between the first feature and the second feature; and a determination unit, used to determine a first target region matching the first image in the second image based on the first target correlation coefficient.

[0024] According to one embodiment of the present invention, a non-volatile storage medium is further provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the image matching method of the embodiment of the present invention when executed by a processor.

[0025] According to one embodiment of the present invention, a processor is further provided, and the processor is used to run a program, wherein the program is configured to execute any of the above-mentioned image matching methods when running.

[0026] According to one embodiment of the present invention, there is further provided an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the image matching method of the embodiment of the present invention.

[0027] In at least some embodiments of the present invention, a first image and a second image are obtained, wherein the first image is an image to be matched, and the second image is a reference image corresponding to the image to be matched; a first feature is extracted from the first image, and a second feature is extracted from the second image; a first target correlation coefficient between the first feature and the second feature is obtained, wherein the first target correlation coefficient is used to indicate the degree of correlation between the first feature and the second feature; and a first target region matching the first image is determined in the second image based on the first target correlation coefficient. That is, the present application determines the first target correlation coefficient using the features of the image to be matched and the features of the reference image corresponding to the image to be matched, and then finds the region matching the given image to be matched in the reference image corresponding to the image to be matched based on the first target correlation coefficient. This method meets the universality conditions in automated testing scenarios, avoids the need for image matching methods based on supervised learning to pre-produce a certain scale of data sets, and is difficult to produce a universal model, thereby solving the technical problem of low efficiency of image matching and achieving the technical effect of improving the efficiency of image matching. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0029] Figure 1 is a hardware structure block diagram of a mobile terminal of an image matching method according to an embodiment of the present invention;

[0030] Figure 2 is a flow chart of an image matching method according to one embodiment of the present invention;

[0031] Figure 3 is a schematic diagram of template matching according to one embodiment of the present invention;

[0032] Figure 4 is a flow chart of an image matching method according to one embodiment of the present invention;

[0033] Figure 5 is a schematic diagram of an OCR-based text matching effect according to one embodiment of the present invention;

[0034] Figure 6 is a schematic diagram of determining a mutual correlation coefficient according to one embodiment of the present invention;

[0035] Figure 7 is a schematic diagram of a characteristic diagram according to one embodiment of the present invention;

[0036] Figure 8 is a schematic diagram of a template image and a bounding box of an original image according to one embodiment of the present invention;

[0037] Fig. 9 is a schematic diagram of post-processing a boundary box of an original image according to one embodiment of the present invention;

[0038] Fig.10 is a schematic diagram of a result of screening multiple bounding boxes according to one embodiment of the present invention;

[0039] Fig.11 is another schematic diagram of a result of screening multiple bounding boxes according to one embodiment of the present invention;

[0040] Fig.12 is a structural block diagram of an image matching device according to one embodiment of the present invention. DETAILED DESCRIPTION

[0041] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0042] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0043] According to one embodiment of the present invention, an embodiment of an image matching method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0044] The method embodiment can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, the mobile terminal can be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID for short), a PAD, a game console and other terminal devices. Figure 1 FIG. 1 is a hardware structure block diagram of a mobile terminal of an image matching method according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1Only one is shown in the figure) processor 102 (processor 102 may include but is not limited to a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processing (DSP) chip, a microprocessor (MCU), a programmable logic device (FPGA), a neural network processor (NPU), a tensor processor (TPU), an artificial intelligence (AI) type processor, etc.) and a memory 104 for storing data. Optionally, the mobile terminal may also include a transmission device 106 for communication functions, an input and output device 108, and a display device 110. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.

[0045] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the image matching method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above-mentioned image matching method is realized. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the mobile terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0046] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0047] The inputs in the input / output device 108 may come from a plurality of human interface devices (HIDs), such as keyboards and mice, game controllers, and other dedicated game controllers (such as steering wheels, fishing rods, dance mats, remote controls, etc.). In addition to providing input functions, some human interface devices may also provide output functions, such as force feedback and vibration of game controllers, audio output of controllers, etc.

[0048] The display device 110 may be, for example, a head-up display (HUD), a touch-screen liquid crystal display (LCD), and a touch display (also referred to as a "touch screen" or "touch display screen"). The liquid crystal display may enable a user to interact with a user interface of the mobile terminal. In some embodiments, the mobile terminal has a graphical user interface (GUI), and a user may interact with the GUI by finger contacts and / or gestures on a touch-sensitive surface, wherein the human-computer interaction functions here may optionally include the following interactions: creating web pages, drawing, word processing, making electronic documents, games, video conferencing, instant messaging, sending and receiving emails, call interfaces, playing digital videos, playing digital music, and / or web browsing, etc. The executable instructions for executing the above-mentioned human-computer interaction functions are configured / stored in a computer program product or a readable storage medium executable by one or more processors.

[0049] In a possible implementation manner, an embodiment of the present invention provides an image matching method. Figure 2 FIG. 1 is a flow chart of an image matching method according to one embodiment of the present invention. Figure 2 As shown, the method may include the following steps:

[0050] Step S202: Acquire a first image and a second image, wherein the first image is the image to be matched, and the second image is a reference image corresponding to the image to be matched.

[0051] In the technical solution provided in the above step S202 of the present invention, a first image may be input, and the first image may be an image to be matched, for example, a template image to be matched given in template matching. This embodiment also inputs a second image, and the second image may be a reference image corresponding to the image to be matched, for example, an original image in template matching. The area of ​​the first image in this embodiment to be matched with it is determined in the second image.

[0052] Step S204: extracting a first feature from the first image and extracting a second feature from the second image.

[0053] In the technical solution provided in the above step S204 of the present invention, after acquiring the first image and the second image, the first feature can be extracted from the first image, and the second feature can be extracted from the second image.

[0054] In this embodiment, a feature extraction operation may be performed on the first image to extract a first feature from the first image. The first feature may also be referred to as a template feature, which may be a feature map. The method of extracting the first feature from the first image may eliminate the influence of background changes in the first image. This embodiment may also perform a feature extraction operation on the second image to extract a second feature from the second image. The second feature may be referred to as a source feature, which may be a feature map. The method of extracting the second feature from the second image may eliminate the influence of background changes in the second image.

[0055] Step S206: Obtain a first target correlation coefficient between the first feature and the second feature, wherein the first target correlation coefficient is used to indicate the degree of correlation between the first feature and the second feature.

[0056] In the technical solution provided in the above step S206 of the present invention, after the first feature is extracted from the first image and the second feature is extracted from the second image, a first target correlation coefficient between the first feature and the second feature can be obtained, and the first target correlation coefficient is used to represent the degree of correlation between the first feature and the second feature.

[0057] In this embodiment, feature matching can be performed on the first feature and the second feature, and a first target correlation coefficient between the first feature and the second feature can be obtained. The first target correlation coefficient can also be called a mutual correlation coefficient or a correlation degree, which is used to indicate the degree of correlation between the first feature and the second feature. The first feature and the second feature can be calculated based on a fast Fourier transform (FFT). For example, the first feature and the second feature can be converted from the time domain to the frequency domain based on a fast Fourier transform to calculate the first target correlation coefficient between the first feature and the second feature. This can greatly improve the processing speed while maintaining the processing accuracy.

[0058] Step S208: determining a first target region matching the first image in the second image based on the first target correlation coefficient.

[0059] In the technical solution provided in the above step S208 of the present invention, after obtaining the first target correlation coefficient between the first feature and the second feature, a first target area matching the first image can be determined in the second image based on the first target correlation coefficient. The first target area can be a small area in the second image that matches the first image, and can be a region of interest (abbreviated as ROI).

[0060] Optionally, this embodiment may output the position information of the boundary box and the center point of the first target area, and the position information may be the coordinate information of the center point.

[0061] The above method of this embodiment can be applied to automated testing scenarios in computer vision, for example, in a one-machine-multiple-control automated testing scenario to achieve target detection and tracking. Among them, the one-machine-multiple-control automated testing scenario can refer to operating one terminal device to record an object, and after recording the object, replaying the recorded object on other N terminal devices.

[0062] Through the above steps S202 to S208, a first image and a second image are obtained, wherein the first image is an image to be matched, and the second image is a reference image corresponding to the image to be matched; a first feature is extracted from the first image, and a second feature is extracted from the second image; a first target correlation coefficient between the first feature and the second feature is obtained, wherein the first target correlation coefficient is used to indicate the degree of correlation between the first feature and the second feature; and a first target region matching the first image is determined in the second image based on the first target correlation coefficient. That is, this embodiment determines the first target correlation coefficient using the features of the image to be matched and the features of the reference image corresponding to the image to be matched, and then finds the region matching the given image to be matched in the reference image corresponding to the image to be matched based on the first target correlation coefficient. This method meets the universality conditions in automated testing scenarios, avoids the need for image matching methods based on supervised learning to pre-produce a certain scale of data sets, and is difficult to produce a universal model, thereby solving the technical problem of low efficiency of image matching and achieving the technical effect of improving the efficiency of image matching.

[0063] The above method of this embodiment is further introduced below.

[0064] As an optional implementation, step S204, extracting a first feature from the first image, includes: extracting a first target feature map from the first image based on a feature extraction model, wherein the feature extraction model is obtained based on convolutional neural network training, and the first target feature map includes: shape features of the first image.

[0065] In this embodiment, when extracting the first feature from the first image, a feature extraction model may be called, the first image may be input into the feature extraction model, and the first image may be processed by the feature extraction model to obtain a first target feature map, and the first feature may include the first target feature map. Optionally, the feature extraction model of this embodiment may be a deep feature extractor, which may be a pre-trained model having shape bias and obtained by training using an unsupervised method, so that the first target feature map may include the shape features of the first image to eliminate the influence of background changes, and it may be obtained by training based on a convolutional neural network (CNN), so that the feature extraction model of this embodiment may be a pre-trained CNN model.

[0066] Optionally, the embodiment may perform mixed training on a stylized dataset (Stylized-ImageNet, referred to as SIN) and a dataset of the largest image recognition database (ImageNet, referred to as IN) to obtain the above-mentioned feature extraction model based on shape representation, that is, the feature extraction model adds an image shape bias. Optionally, the feature extraction model may be a visual geometry group (Visual Geometry Group, referred to as VGG), which is a new deep convolutional neural network model, such as VGG-19, which may also be referred to as VGG19_SIN_IN (mixed training by SIN&IN datasets), so as to extract a first target feature map from the first image.

[0067] In related technologies, convolutional neural networks tend to use color and texture for prediction, but this is different from the way humans identify objects by shape. In this embodiment, a feature extraction model with an image shape bias is used to extract a first target feature map from the first image, so that the shape features in the first image can be better extracted, so that the image content can be more accurately portrayed.

[0068] As an optional implementation, the method further includes: outputting a first target feature map based on a feature output layer of the feature extraction model, wherein the feature output layer is determined based on a size of the first image.

[0069] In this embodiment, the feature extraction model may include a feature output layer, which may be determined based on the size of the first image, thereby achieving adaptive adjustment of the feature output layer and achieving the purpose of scale adaptation, wherein the size of the first image may be a CNN feature output layer, thereby outputting a first target feature map based on it and using it as input for a subsequent feature matching algorithm.

[0070] The difference between the size of the first image and the size of the second image may be relatively large. In the related art, the size of the first image and the size of the second image are usually resized and then used as CNN input of the same size, which may lead to information loss and feature matching failure. The feature output layer of the feature extraction model of this embodiment is determined based on the size of the first image, and adaptive adjustment of the feature output layer is achieved, thereby avoiding information loss and feature matching failure.

[0071] As an optional implementation, step S204, extracting a second feature from the second image, includes: extracting a second target feature map from the second image based on a feature extraction model, wherein the second target feature map includes: a shape feature of the second image.

[0072] In this embodiment, when extracting the second feature from the second image, the trained feature extraction model may be called, the second image may be input into the feature extraction model, and the second image may be processed by the feature extraction model to obtain a second target feature map, and the second feature may include the second target feature map. Since the feature extraction model of this embodiment may be a pre-trained model with shape bias, the second target feature map may include the shape features of the second image to eliminate the influence of background changes.

[0073] As an optional implementation, the method further includes: outputting a second target feature map based on a feature output layer of the feature extraction model, wherein the feature output layer is determined based on a size of the first image.

[0074] In this embodiment, the feature output layer of the feature extraction model can be determined based on the size of the first image, thereby realizing adaptive adjustment of the feature output layer and achieving the purpose of scale adaptation. This embodiment can output the second target feature map based on the feature output layer and use it as input for a subsequent feature matching algorithm to avoid information loss and feature matching failure.

[0075] As an optional implementation, step S206, obtaining a first target correlation coefficient between the first feature and the second feature, includes: transforming the first target feature map from the time domain to the frequency domain to obtain a third target feature map; transforming the second target feature map from the time domain to the frequency domain to obtain a fourth target feature map; performing normalized cross-correlation processing on the third target feature map and the fourth target feature map to obtain the first target correlation coefficient.

[0076] In this embodiment, when obtaining the first target correlation coefficient between the first feature and the second feature, the first target feature map can be transformed from the time domain to the frequency domain, for example, the first target feature map in the time domain is subjected to a fast Fourier transform to obtain a third target feature map in the frequency domain. Optionally, this embodiment also transforms the second target feature map from the time domain to the frequency domain to obtain a fourth target feature map, for example, the second target feature map in the time domain is subjected to a fast Fourier transform to obtain a fourth target feature map in the frequency domain. After obtaining the third target feature map and the fourth target feature map in the frequency domain, the third target feature map and the fourth target feature map in the frequency domain can be subjected to a normalized cross-correlation (Normalized Cross-Correlation, referred to as NCC) to obtain the first target correlation coefficient.

[0077] Optionally, if the first image is represented by t, its size may be N x ×N y , the second image is represented by f, and its size can be M x ×M y , the first image and the second image are processed by NCC, which can be a sliding window method (pixel-by-pixel) of the first image on the second image to calculate the correlation coefficient between the first image f and the second image t at each point (u, v), and the correlation coefficient matrix γ can be obtained u,v , can be expressed as the correlation coefficient matrix γ u,v The maximum value of γ max as the best matching position.

[0078]

[0079] Where u∈{0, 1, 2, ..., M x -N x}, v∈{0, 1, 2, ..., M y -N y}. It is used to represent the pixel mean of the first image within the moving area of ​​the second image f(x, y), The definition can be as follows:

[0080]

[0081] However, the above NCC has a very high computational cost, and its computational complexity is N x N y (M x -N x )(M y -N y), and it increases exponentially with the increase of image scale. In addition, when the background of the first image and the second image is cluttered or complex deformation occurs, the matching performance is poor and the calculation amount is huge, which cannot meet the requirements of real-time reasoning of image matching.

[0082] However, in this embodiment, the NCC calculation method based on fast Fourier transform is adopted to perform normalized cross-correlation processing on the third target feature map in the frequency domain corresponding to the first image and the fourth target feature map in the frequency domain corresponding to the second image to obtain the first target correlation coefficient, which can greatly reduce the matching time, and there is no information loss in the conversion between the time domain and the frequency domain, and the matching accuracy can be consistent with the NCC.

[0083] Optionally, in this embodiment, for the above formula (1), it can be equivalently converted from the time domain to the frequency domain for calculation to obtain the Fourier correlation coefficient in the frequency domain as shown in formula (3):

[0084] r(u,v)=∑ x,y f(x, y)·t(xu, y+v)

[0085] R(u,v)=F(u,v)·T(u,v) (3)

[0087] As an optional implementation, normalized cross-correlation processing is performed on the third target feature map and the fourth target feature map to obtain a first target correlation coefficient, including: determining a first complex conjugate value corresponding to the third target feature map and a second complex conjugate value corresponding to the fourth target feature map; performing an inverse Fourier transform on the product of the first complex conjugate value and the second complex conjugate value to obtain the first target correlation coefficient.

[0088] In this embodiment, after the first target feature map is transformed from the time domain to the frequency domain to obtain the third target feature map; after the second target feature map is transformed from the time domain to the frequency domain to obtain the fourth target feature map, it can be known from formula (3) that the third target feature map can be complex conjugated to obtain the first complex conjugate value T(u, v), and the fourth target feature map can be complex conjugated to obtain the second complex conjugate value F(u, v), and then the first complex conjugate value and the second complex conjugate value are multiplied to obtain the product R(u, v)=F(u, v)·T(u, v), and then the inverse Fourier transform (inverse FFT) is performed to obtain the first target correlation coefficient, as shown in the following formula:

[0089]

[0090] In this embodiment, the NCC calculation complexity after FFT transformation is M x My log2(M x M y ), thereby greatly reducing the complexity of correlation coefficient calculation.

[0091] As an optional implementation, the first target feature map is a multi-channel first target feature map, and the second target feature map is a multi-channel second target feature map. Step S206, obtaining the first target correlation coefficient between the first feature and the second feature, includes: obtaining the first target correlation coefficient between the first target feature map of each channel and the second target feature map of each channel, and obtaining multiple first target correlation coefficients; step S208, determining the first target area matching the first image in the second image based on the first target correlation coefficient, includes: determining the first target area in the second image based on the multiple first target correlation coefficients.

[0092] In this embodiment, the first target feature map may be a first target feature map of multiple channels (dimensions), for example, a first target feature map of C channels, which may be obtained by F t It can be represented by a grayscale image, where C can be 512, which is not specifically limited here. The second target feature map of this embodiment can be a multi-channel second target feature map, for example, it can be a second target feature map of C channels, and the second target feature map can be obtained by F f This embodiment can traverse the multi-channel first target feature map and the corresponding multi-channel second target feature map, respectively calculate the first target correlation coefficient between the first target feature map of each channel and the corresponding second target feature map of each channel, so as to obtain multiple first target correlation coefficients, and the correlation coefficient matrix can be determined by the multiple first target correlation coefficients, and the correlation coefficient matrix can be called a cross-correlation matrix, and then the first target area is determined in the second image.

[0093] As an optional implementation, determining the first target area in the second image based on multiple first target correlation coefficients includes: obtaining the maximum first target correlation coefficient among multiple first target correlation coefficients; adjusting the maximum first target correlation coefficient based on a target adjustment parameter; determining a correlation coefficient among the multiple first target correlation coefficients that is greater than or equal to the adjusted maximum first target correlation coefficient as a second target correlation coefficient; and determining the first target area in the second image based on the second target correlation coefficient.

[0094] In this embodiment, the maximum first target correlation coefficient in the correlation coefficient matrix can be obtained, for example, γ max, and then adjust the maximum first target correlation coefficient based on the target adjustment parameter. For example, the target adjustment parameter may be thr, which may be 0.98, which may be a fixed value determined after a large number of tests. This embodiment may multiply the target adjustment parameter and the maximum first target correlation coefficient to obtain the adjusted maximum first target correlation coefficient. This embodiment may determine the correlation coefficient of the multiple first target correlation coefficients that is greater than or equal to the adjusted first target correlation coefficient as the second target correlation coefficient. That is, the second target correlation coefficient may be:

[0095] r u,v ≥thr*r max (5)

[0096] This embodiment can retain all the second target correlation coefficients γ u,v , and then the first target area can be determined in the second image based on all the second target correlation coefficients.

[0097] It should be noted that, due to the difference in background changes between the first image and the second image, if only the maximum value γ in the correlation coefficient matrix is ​​obtained, max , it may cause inaccurate image matching or even image matching failure. However, this embodiment adjusts the maximum first target correlation coefficient based on the target adjustment parameter, and then determines the first target area in the second image based on the adjusted maximum first target correlation coefficient, thereby achieving the purpose of image matching, improving the accuracy of image matching, and ensuring the success rate of image matching.

[0098] As an optional implementation, determining the first target area in the second image based on the second target correlation coefficient includes: determining first position information corresponding to the second target correlation coefficient; determining second position information in the second image based on the first position information; and determining the first target area based on the second position information.

[0099] In this embodiment, when determining the first target area in the second image based on the second target correlation coefficient, the first position information corresponding to the second target correlation coefficient may be determined first, and the first position information may be the coordinate value (u, v) corresponding to the second target correlation coefficient, that is, the output coordinate value (u, v) that satisfies the condition. This embodiment may determine the corresponding second position information in the second image based on the first position information, and the second position information may be the position information of the center of the first target area, and then determine the first target area in the second image based on the second position information.

[0100] As an optional implementation, determining the second position information in the second image based on the first position information includes: converting the first position information into the second position information based on the width and height of the first target feature map, the width and height of the second target feature map, and the scaling ratio of the second image to the second target feature map.

[0101] In this embodiment, when determining the second position information in the second image based on the first position information, the width of the first target feature map may be determined first, which can be obtained by F t,width It can also be expressed as, and the height of the first target feature map can be determined, which can be obtained by F t,hight This embodiment can also determine the width of the second target feature map, which can be expressed by F f,width It can also be expressed as, and the height of the second target feature map can be determined, which can be obtained by F f,hight This embodiment can also obtain the scaling ratio of the second image to the second target feature map, for example, the size of the second image is M x ×M y , then the scaling ratio can be Thus, this embodiment converts the first position information into the second position information based on the width and height of the first target feature map, the width and height of the second target feature map, and the scaling ratio of the second image to the second target feature map.

[0102] Optionally, the first position information of this embodiment may be a coordinate value (u, v), which may be converted into second position information in the second image. The second position information may be the coordinates (x center ,y center ), the first position information can be converted into the second position information based on the width and height of the first target feature map, the width and height of the second target feature map, and the scaling ratio of the second image to the second target feature map by the following formula:

[0103]

[0104] As an optional implementation, determining the first target area based on the second position information includes: determining the second position information as the position information of the center of the first target area; and determining a bounding box of the first target area based on the position information of the center to obtain the first target area.

[0105] In this embodiment, when determining the first target area based on the second position information, the second position information may be determined as the position information of the center of the first target area. For example, the position of the center may be the coordinates (x center ,y center), and then a bounding box can be generated based on it, and then the first target area can be determined in the second image through the position information of the bounding box and the center.

[0106] As an optional implementation, in the case where the number of second target correlation coefficients is multiple, the number of bounding boxes is multiple, and the bounding box of the first target area is determined based on the position information of the center to obtain the first target area, including: selecting a target bounding box from the multiple bounding boxes based on the intersection-and-union ratio between the multiple bounding boxes; determining that the number of target bounding boxes is multiple, then selecting a first target bounding box from the multiple target bounding boxes based on the target points in the second image; and determining the area of ​​the first target bounding box in the second image as the first target area.

[0107] In this embodiment, since the second target correlation coefficient r u,v ≥thr*r max , the number of which can be multiple, and thus the number of its corresponding coordinate values ​​(u, v) can also be multiple, and the position information of the center of the first target area determined by it can also be multiple, to generate multiple bounding boxes, which can be multiple redundant bounding boxes at the same position on the second image. This embodiment can select a target bounding box from multiple bounding boxes based on the intersection-and-union ratio between multiple bounding boxes. Optionally, this embodiment can use non-maximum suppression (NMS) to select a target bounding box from multiple bounding boxes based on the intersection-and-union ratio between multiple bounding boxes, and the purpose can be to eliminate a large number of redundant bounding boxes at the same position. In this embodiment, when using NMS to screen from multiple bounding boxes (candidate boxes), it can be based on γ within a certain range within the bounding box. u,v The sum is used as the sorting basis for sorting multiple bounding boxes, and the intersection over union (IoU) of the sorted multiple bounding boxes is greater than the target value as the elimination criterion. For example, the target value can be 0.5. Multiple bounding boxes are repeatedly eliminated until the bounding box list is empty, and the target bounding box is output.

[0108] Optionally, in this embodiment, after selecting the target bounding box from multiple bounding boxes, multiple target bounding boxes (maximum boxes, multi-target problems) may still be retained, and then a unique first target bounding box may be further selected from the multiple target bounding boxes based on the target point in the second image. Optionally, this embodiment may select a unique bounding box from multiple target bounding boxes by a reference point coordinate suppression method, and then output the unique bounding box, and use it to determine the first target area in the second image, wherein the reference point coordinate suppression method is to use the screen coordinate point clicked by the user as the reference point, and discard the algorithm output result outside the normalized distance determination range. Optionally, the target point in the second image of this embodiment can be a coordinate point obtained by performing a touch operation on the screen of the device, which can be used as a reference point A, and the corresponding center point output by the model is B, which can meet the precondition that the center point coordinates of the real matching area of ​​multiple devices are not far from the reference point coordinates when one machine is multi-controlled by default. When A and B satisfy the following formula, that is, the normalized distance is less than 0.1, then B can be output:

[0109] norm(|A i -B i |)<0.1 (7)

[0110] Among them, i is used to represent the i-th reference point A and the corresponding midline point B.

[0111] As an optional implementation, the receptive field size of the network layer of the feature extraction model does not exceed the size of the second image.

[0112] In this embodiment, the feature extraction model requires a certain receptive field size limit. For example, if the feature extraction model is CNN, the performance can be guaranteed by limiting the size of the receptive field. Optionally, the receptive field size RF of the network layer of the feature extraction model of this embodiment does not exceed the size S of the second image, as shown in the following formula:

[0113] RF≤S (8)

[0114] Among them, the calculation of the receptive field size RF of the feature map (first target feature map or second target feature map) output by the i-th layer network of the feature extraction model can be shown as follows:

[0115] RF i =RF i-1 +(k-1)j i-1 (9)

[0116] Among them, k can be used to represent the kernel size of the i-th layer of the feature extraction model, and j can be used to represent the interval between the features of the output feature map, which is equal to the interval value of the previous layer multiplied by the step size of the convolution.

[0117] This embodiment can calculate the size of the receptive field of each layer of the feature extraction model by the above formula (9). Optionally, this embodiment can default that the size of the first image can be 1 / 10 of the size of the second image (the size can be adjusted, but basically will not change drastically), which can be in the range of 216 to 256. Therefore, after adaptive feature selection, this embodiment can mostly use the conv5_n layer network as the feature output layer. The deeper the number of layers, the more satisfying the prerequisite for extracting rich semantic information.

[0118] As an optional implementation, the method also includes: if it is determined that the first image includes first text information, the first text information is extracted from the first image, and the second text information is extracted from the second image; fuzzy matching is performed on the first text information and the second text information; if it is determined that the fuzzy matching of the first text information and the second text information is successful, a second target area matching the first text information is determined in the second image.

[0119] In this embodiment, it can be determined whether the first image includes the first text information. Optionally, this embodiment can perform optical character recognition (OCR) text recognition on the first image to determine whether the first image includes the first text information. If the first image includes the first text information, for example, the first image is a text template image including template words, then OCR recognition can be performed on the second image to obtain the second text information. For example, the second text information can be the original image text (source words). Optionally, this embodiment can perform text detection (DBnet) and text recognition (CRNN) on the first image to infer the first text information contained in the first image and the second text information contained in the second image. After the first text information is extracted from the first image and the second text information is extracted from the second image, text fuzzy matching can be performed on the first text information and the second text information. If the fuzzy matching of the first text information and the second text information is successful, a second target area matching the first text information can be determined in the second image.

[0120] As an optional implementation, determining the second target area matching the first text information in the second image includes: determining third position information of the first text information in the second image; and determining the second target area based on the third position information.

[0121] In this embodiment, when determining the second target area matching the first text information in the second image, the third position information of the first text information in the second image can be determined first, and then the second target area can be determined based on the third position information. The position information of the text detection box and the center point can be determined in the second image based on the third position information, and then the position information of the text detection box and the center point can be output.

[0122] As an optional implementation, step S204, extracting the first feature from the first image and extracting the second feature from the second image, includes: determining that the first image does not include the first text information, or determining that fuzzy matching of the first text information and the second text information fails, then extracting the first feature from the first image and extracting the second feature from the second image.

[0123] In this embodiment, if it is determined by performing OCR recognition on the first image that the first text information is not included in the first image, the first feature can be extracted from the first image, and the second feature can be extracted from the second image. Alternatively, if it is determined that fuzzy matching of the first text information and the second text information fails, the first feature can be extracted from the first image, and the second feature can be extracted from the second image, so as to determine a first target area matching the first image in the second image based on a first target correlation coefficient between the first feature and the second feature.

[0124] As an optional implementation, the method further includes: obtaining the similarity between the image corresponding to the first target area and the first image; and if it is determined that the similarity is greater than a target threshold, outputting a prompt message, wherein the prompt message is used to indicate that the first image and the second image are successfully matched.

[0125] In this embodiment, after determining the first target area matching the first image in the second image based on the first target correlation coefficient, the similarity between the image corresponding to the first target area and the first image can be calculated, and then it is determined whether the similarity is greater than a target threshold. The target threshold is a critical threshold for measuring the degree of similarity between the image corresponding to the first target area and the first image, and can be a threshold of cosine similarity. The calculation result of the similarity between the image corresponding to the first target area and the first image in this embodiment is based on a certain target threshold. For example, the target threshold is 0.9. If the similarity is greater than 0.9, it indicates that the image corresponding to the first target area and the first image are similar. Otherwise, it indicates that the image corresponding to the first target area and the first image are not similar.

[0126] If it is determined that the similarity is greater than the target threshold, prompt information may be output, and it is determined that the first image matches the second image successfully, and a first target area matched by the first image in the second image may be output.

[0127] Optionally, when the first text information exists in the first image, the similarity between the image corresponding to the second target area and the first image can be calculated, and then it can be determined whether the similarity is greater than a target threshold. If it is determined that the similarity is greater than the target threshold, a prompt message can also be output to determine that the first image and the second image are successfully matched, and the second target area matched by the first image in the second image can be output.

[0128] Optionally, if it is determined that the similarity is not greater than the target threshold, a null value (None) can be output for manual intervention. For example, when testing a game application in a one-machine multi-control automated testing scenario, it is now necessary to select map A, where one of the terminal devices will stop after the image matching fails, and the user is required to manually click on map A for manual intervention.

[0129] Optionally, if the B map is entered by mistake, a false positive problem occurs, and the tester needs to exit the game application and reselect the map, etc. To avoid the false positive problem, this embodiment can calculate the similarity between the image corresponding to the first target area and the first image, and output None when the similarity is not greater than the target threshold, and then perform manual intervention.

[0130] Optionally, in this embodiment, when calculating the similarity between the image corresponding to the first target area and the first image, the feature F of the image corresponding to the first target area may be input. b and the first target feature map F of the first image t , then the similarity result confidence can be calculated by the following formula:

[0131]

[0132] Among them, flatten is used to represent the F of multiple channels (for example, C channel) b and F t Flatten to one dimension.

[0133] This embodiment uses a pre-trained CNN model with scale adaptation and shape bias as a feature extraction module, combined with a fast NCC based on fast Fourier transform as the core matching algorithm, which still has strong robustness under complex conditions such as background changes. In the case where there is text in the first image, OCR text recognition can be introduced for character fuzzy matching. In addition, in order to avoid false successes during the automated test process, post-processing solutions such as non-maximum suppression, reference point coordinate suppression, and similarity calculation are added to produce an image matching system with high robustness, real-time reasoning, and the ability to effectively avoid false successes.

[0134] The technical solution of this embodiment is further illustrated below by taking the first image as a template image and the second image as an original image as an example.

[0135] In this embodiment, the template matching algorithm is one of the basic tasks in computer vision and can be applied to target detection, tracking and other fields. Figure 3 FIG. 1 is a schematic diagram of template matching according to one embodiment of the present invention. Figure 3 As shown, a small area position c that matches the given template image a can be found in the original image b, and the bold frame is the matched small area position c. In automated testing, template matching algorithms also play an important role in locating regions of interest.

[0136] In the related art, template matching algorithms may include the following methods: represented by traditional operators such as object recognition algorithms (SIFT / SURF), which can match the number of local invariant feature points, and use algorithms such as random sample consensus (RANSAC) and pattern matching (BF) to eliminate mismatched points; (2) the template image and the sub-window of the original image can be measured pixel by pixel similarity by sliding window, and the three-channel image can be converted into a grayscale image for calculation, using the sum of absolute differences (SAD), statistics and data analysis (CSAD), NCC and other algorithms as similarity measurement methods; (3) deep learning can be used as a local image matching solution, and the mainstream architecture is a network for training and solving similarity functions (Siamese) and a triplet abstract data (triplet) network. For example, it can be a dual-branch weight sharing network (MatchNet), a local block descriptor (L2-Net) and other algorithms, which can be divided into two categories: with a metric layer and without a metric layer.

[0137] Due to the need to meet the conditions of universality in automated scenarios, related technologies still have various inadaptabilities. For example, the local feature point matching scheme represented by SIFT / SURF in the above method (1) relies too much on prior knowledge, resulting in poor robustness of image matching in different scenarios, especially in the case of scale changes or smooth areas of the template or original image, the image matching performance drops sharply; the similarity measurement method based on pixel-by-pixel sliding window in the above method (2) performs poor performance when the background of the template or original image is cluttered, or when the template or original image undergoes complex deformation, and the amount of calculation is huge, which cannot meet the requirements of real-time reasoning of image matching; the image matching scheme based on supervised learning in the above method (3) can have a much higher matching accuracy than the above schemes, but because it requires a certain size of data set for training, it requires a lot of manpower and time costs in the early stage, and its generalization is also difficult to predict.

[0138] In addition, in the scenario of one-machine-multiple-control automated testing, the difficulties of image matching mainly focus on image geometry changes, template background texture changes, multiple smooth areas, false successful matches, real-time reasoning and other issues; the image matching algorithm based on supervised learning is more robust in different scenarios, but requires a large number of training sets to be created in advance and has generalization problems, making it difficult to produce a general model; because different matching algorithms will eventually output the value with the highest confidence, but it is not necessarily the actual optimal solution, if a false success problem occurs in automated testing, it will cause a huge reset / rollback burden on testers.

[0139] Therefore, in response to the above problems, this embodiment uses a pre-trained CNN model with scale adaptation characteristics and shape bias as a deep feature extractor, combined with NCC based on fast Fourier transform FFT as the core matching algorithm, which still has strong robustness under complex conditions such as background changes. In the case where the foreground of the template image is text, OCR text recognition is introduced for character fuzzy matching. In addition, in order to avoid false successes during the automated test process, post-processing solutions such as non-maximum suppression, reference point coordinate suppression and similarity calculation can be added to produce an image matching system with high robustness, real-time reasoning and the ability to effectively avoid false successes.

[0140] The above method of this embodiment is further introduced below.

[0141] Figure 4 FIG. 1 is a flow chart of an image matching method according to one embodiment of the present invention. Figure 4 As shown, the method may include the following steps:

[0142] Step S401, obtaining a template image.

[0143] Step S402, obtaining the original image.

[0144] Step S403: perform OCR recognition on the template image.

[0145] Step S404, entering the process of performing OCR text matching on the template image and the original image.

[0146] This embodiment can perform OCR recognition on the template image, extract text information from the input template image, and then enter the process of performing OCR text matching on the template image and the original image.

[0147] Step S405: extracting the template text from the template image.

[0148] Step S406, extracting the original image text from the original image.

[0149] Step S407: perform text fuzzy matching on the template image text and the original image text.

[0150] Step S408: output the detection box and center point coordinates of the matched text.

[0151] In this embodiment, when the text fuzzy matching between the template image text and the original image text is successful, the text boundary box and center point coordinates corresponding to the template image in the original image can be directly output.

[0152] Step S409, enter the process of performing CNN feature matching between the template image and the original image.

[0153] This embodiment can enter the process of CNN feature matching of the template image and the original image when performing OCR recognition on the template image and determining that there is no text information in the template image, or when determining that text fuzzy matching of the template image text and the original image text fails.

[0154] Step S410: extracting template features from the template image.

[0155] Step S411, extracting original image features from the original image.

[0156] Step S412: Calculate the correlation coefficient between the template image features and the original image features based on FFT.

[0157] Step S413: determining the area corresponding to the template image in the original image based on the mutual correlation coefficient.

[0158] This embodiment determines the area corresponding to the template image in the original image based on the mutual correlation coefficient, and outputs a unique bounding box after post-processing multiple bounding boxes corresponding to the area, which is then displayed in the original image.

[0159] This embodiment can use CNN feature matching to determine the final ROI area and the position of its center point.

[0160] The template matching algorithm based on OCR text recognition of this embodiment is introduced below.

[0161] In this embodiment, in the OCR-based template matching method, the classic two-stage algorithm of text detection (DBnet) + text recognition (CRNN) can be used to infer text information for fuzzy matching, determine the corresponding position of the text in the template image in the original image, and then output the coordinates of the text detection box and the center point coordinates.

[0162] For example, for scenes with a lot of artistic fonts in mobile games, targeted training samples can be created to fine-tune the OCR recognition model to improve the algorithm recognition effect. Figure 5 As shown, for example, the template image is a template image including the word "enter", and "enter" can be identified in the original image and the template image to determine the area d matching the template image including the word "enter" in the original image. Figure 5 It is a schematic diagram of an OCR-based text matching effect according to one embodiment of the present invention.

[0163] It should be noted that the OCR-based text matching effect in the above embodiment is only an example of the embodiment of the present invention, and does not limit the OCR-based text matching effect in the embodiment of the present invention.

[0164] The following is an introduction to the unsupervised template matching algorithm based on pre-trained CNN of this embodiment.

[0165] In this embodiment, CNN feature matching can be divided into two modules: feature extraction and correlation calculation.

[0166] In the feature extraction part, a pre-trained model with shape bias can be used to eliminate the impact of background changes, and the feature output layer can be adaptively adjusted to ensure the same deep feature representation (template image and original image);

[0167] In the correlation calculation part, the feature map of the template image and the feature map of the original image can be converted from the time domain to the frequency domain based on the fast Fourier transform to calculate the mutual correlation coefficient, which greatly improves the inference speed during image matching while maintaining accuracy.

[0168] The feature extraction part of this embodiment is introduced below.

[0169] In this embodiment, the Stylized-ImageNet (SIN) and ImageNet (IN) data sets can be mixed and trained to obtain a pre-trained model VGG19_SIN_IN based on shape representation, that is, a pre-trained model CNN with shape bias, which is used to achieve feature extraction. In related technologies, convolutional neural networks tend to use color and texture for prediction, which is different from the way humans distinguish objects by shape. In this embodiment, the VGG19_SIN_IN model obtained by mixed training of the SIN&IN data sets can be used to increase the image shape bias, and better shape features can be extracted from the template image and the original image to more accurately characterize the content of the template image and the original image.

[0170] In this embodiment, in template matching, the size of the template image is usually quite different from the size of the original image. In the related art, after the template image and the original image are resized, they can be used as CNN input with a uniform size. This method will cause information loss and feature matching failure. However, in this embodiment, the feature output layer of the CNN can be determined based on the size of the template image, thereby achieving input scale adaptation, outputting feature maps of the template image and feature maps of the original image, and using them as inputs for subsequent feature matching algorithms.

[0171] For CNN, a certain receptive field size limit is required to ensure performance. Optionally, the receptive field size of the highest layer of CNN should not exceed the size S of the original image:

[0172] RF≤S (11)

[0173] Among them, the receptive field size RF of the i-th layer network output feature map of CNN is calculated as follows:

[0174] RF i =RF i-1 +(k-1)j i-1 (12)

[0175] Among them, k can be used to represent the kernel size of the i-th layer of CNN; j can be used to represent the interval between the features of the output feature map, which is equal to the interval value of the previous layer multiplied by the step size of the convolution.

[0176] This embodiment can calculate the receptive field size of each layer of CNN by the above formula (12). In this embodiment, the size of the template image can be assumed to be 1 / 10 of the original image size (the size can be adjusted, but there will be basically no drastic changes), which can be in the range of 216 to 256. Therefore, this embodiment can use the convolutional layer conv5_n layer network as the feature output layer after adaptive feature selection. The deeper the number of layers, the more satisfying the prerequisite for extracting rich semantic information.

[0177] The correlation calculation part of this embodiment is further introduced below.

[0178] NCC is an image matching method. For example, the template image in the two images to be matched is t and the size is N. x ×N y , the original image can be f, and the size can be M x ×M y , which uses the template image to slide on the original image in a pixel-by-pixel manner to calculate the correlation coefficient between f and t at each point (u, v) to obtain the correlation coefficient matrix γ u,v , and with γ u,v The maximum value of γ max as the best matching position.

[0179]

[0180] Where u∈{0, 1, 2, ..., M x -N x}, v∈{0, 1, 2, ..., M y -N y}. It can be used to represent the pixel mean of the template image in the moving area of ​​the original image f(x, y). The definition can be as follows:

[0181]

[0182] However, the above NCC has a very high computational cost, and its computational complexity is N x N y (M x -N x )(M y -N y ), and increases exponentially with the increase of the scale of the template image and the original image.

[0183] In this embodiment, the NCC calculation method based on fast Fourier transform can be used. After testing, the time length of image matching can be greatly reduced, and there is no information loss in the mutual conversion between time domain and frequency domain in theory, and the accuracy test is consistent with NCC. For the above formula (13), the time domain can be equivalently converted to the frequency domain for calculation, such as the following formula (15), and the Fourier correlation coefficient in the frequency domain can be obtained:

[0184] r(u,v)=∑x,yf(x,y)·t(xu,y+v)

[0185] R(u,v)=F(u,v)·T(u,v) (15)

[0187] Figure 6 FIG. 1 is a schematic diagram of determining a correlation coefficient according to one embodiment of the present invention. Figure 6 As shown, the feature map of the template image can be fast Fourier transformed to transfer it from the time domain to the frequency domain. The feature map of the original image can be fast Fourier transformed to transfer it from the time domain to the frequency domain. Then, the complex conjugate values ​​of the feature map in the frequency domain of the template image and the feature map in the frequency domain of the original image are calculated to obtain T(u, v) and F(u, v). Then, they are multiplied and the obtained product R(u, v) = F(u, v) · T(u, v) is inverse Fourier transformed, as shown in the following formula, and finally the mutual correlation coefficient is obtained:

[0188]

[0189] In this embodiment, the feature map of the template image is transformed from the time domain to the frequency domain by performing a fast Fourier transform, the feature map of the original image is transformed from the time domain to the frequency domain by performing a fast Fourier transform, and then the NCC calculation is performed. The complexity is M x M y log2(M x M y ), thus this embodiment greatly reduces the complexity of calculation by obtaining the cross-correlation coefficient through FFT-based NCC.

[0190] The correlation coefficient matrix of this embodiment is further introduced below.

[0191] In this embodiment, the template image and the original image output feature maps of C channels, which can be denoted as F t With F f , then traverse the C channels and calculate the mutual correlation coefficients between the corresponding feature maps respectively, where the feature image of each channel can be regarded as a grayscale image, and finally accumulate the results to obtain the final correlation coefficient matrix. Figure 7 is a schematic diagram of a characteristic diagram according to one embodiment of the present invention. Figure 7As shown, the left image can be the original image, the middle image is the visualization result of the feature map of a single channel, and the right image can be the visualization result of the mean result of the feature map of 512 channels, which has stronger expressiveness.

[0192] Due to the differences in background changes, if only the maximum value γ in the correlation coefficient matrix is ​​obtained max , which may lead to inaccurate image matching or even matching failure. Therefore, this embodiment can set a threshold, which can be a fixed value determined after a large number of tests. When thr = 0.98 and satisfies formula (17), all γ that meet the conditions can be retained. u,v , and its corresponding coordinate values ​​(u,v).

[0193] r u,v ≥thr*r max (17)

[0194] The method for generating the center point coordinates and the bounding box of this embodiment is further described below.

[0195] In this embodiment, the output coordinate value (u, v) that meets the conditions can be converted into the center point coordinate (x center ,y center ), obtained by the following formula:

[0196]

[0197] Among them, F t,width 、F t,hight It can be used to represent the width and height in the feature map of the template image, F f,width 、F f,hight It can be used to represent the width and height in the feature map of the original image. It can be used to represent the scaling ratio of the feature map from the original image to the original image.

[0198] Figure 8 FIG. 1 is a schematic diagram of a template image and a boundary box of an original image according to one embodiment of the present invention. Figure 8 As shown in the figure, the left side is the template image e to be matched, and the line boxes in the original image on the right are all the bounding boxes corresponding to the template image e generated according to the center point coordinates. u,v The choice exists thr*γ max The number of the center point in the original image can be multiple, and the corresponding coordinate value (u, v) can be multiple. Therefore, the center point coordinates in the original image can be multiple, and the number of the corresponding bounding boxes can also be multiple.

[0199] In order to avoid the serious impact of false success on the automated test process, this embodiment can add a variety of post-processing schemes, for example, through non-maximum suppression, the redundant bounding boxes of multiple bounding boxes at the same position can be eliminated; based on the reference point coordinate suppression, the user clicks the screen coordinate point as the reference point, and the algorithm output results outside the normalized distance determination range are discarded; the similarity between the template image and the ROI area of ​​the original image can be calculated, and only the output results greater than a given threshold are retained. It is further introduced below.

[0200] Fig. 9 FIG. 1 is a schematic diagram of post-processing a boundary box of an original image according to one embodiment of the present invention. Fig. 9 As shown, in this embodiment, after the template image and the original image are matched based on the CNN algorithm, the matching result can be further post-processed. Non-maximum suppression (NMS) is a commonly used bounding box post-processing method in image processing fields such as target detection. As an important component of algorithms such as YOLO, faster rcnn, and SSD, its purpose is to eliminate a large number of redundant bounding boxes at the same position. Fig.10 As shown, NMS can be used to filter multiple bounding boxes, where the bounding boxes within a certain range of γ u,v The sum is used as the sorting basis, and the intersection over union (IoU) of the sorted bounding boxes is greater than 0.5 as the elimination criterion. The elimination process is repeated until the bounding box list is empty, and the final result is output. Fig.10 is a schematic diagram of a result of screening multiple bounding boxes according to one embodiment of the present invention. After the multiple bounding boxes are screened by NMS, the number of bounding boxes can be reduced to bounding box f and bounding box g.

[0201] In this embodiment, in some cases, if multiple bounding boxes (multiple maximum value boxes, multi-target problems) are still retained after NMS processing, a unique bounding box will be further output from the filtered bounding boxes through the reference point coordinate suppression process. Optionally, this embodiment allows the user to click the device screen coordinates as the reference point A, and the model outputs the center point as A. It defaults to the premise that the coordinates of the center points of the real matching areas of multiple devices are not far from the coordinates of the reference points when one machine has multiple controls. When A and B satisfy the following formula, that is, the normalized distance is less than 0.1, B can be output:

[0202] norm(|A i -B i |)<0.1 (19)

[0203] Fig.11 FIG. 1 is another schematic diagram of the result of screening multiple bounding boxes according to one embodiment of the present invention. Fig.11As shown, a unique bounding box g is output from the filtered bounding boxes f and g through the reference point coordinate suppression process.

[0204] Optionally, if a bounding box is still retained after the NMS process, the reference point coordinate suppression process may not be performed on the bounding box.

[0205] In this embodiment, in order to avoid the problem of false success, the template image and the image corresponding to the target area of ​​the original image can be similarly calculated, and when the calculation result is not greater than a given threshold, None is output and manual intervention is performed. When the calculation result is greater than the given threshold, the area R in the original image that matches the template image can be output.

[0206] Optionally, in this embodiment, the input of the similarity calculation may be the feature F of the template image. t The feature F of the image corresponding to the target area b , then the similarity result confidence can be calculated by the following formula:

[0207]

[0208] Among them, flatten can be used to indicate flattening the C channel feature map to one dimension.

[0209] This embodiment uses a pre-trained CNN model with scale adaptation characteristics and shape bias as a deep feature extractor, combined with NCC based on fast Fourier transform FFT as the core matching algorithm, which still has strong robustness under complex conditions such as background changes. In the case where the foreground of the template image is text, OCR text recognition is introduced for character fuzzy matching. In addition, in order to avoid false successes during the automated test process, post-processing solutions such as non-maximum suppression, reference point coordinate suppression and similarity calculation can be added to produce an image matching system with high robustness, real-time reasoning and the ability to effectively avoid false successes, and realize a template matching algorithm with universality, high precision and real-time reasoning. When performing unified automated tests on dozens of mobile phones, it can effectively improve the image matching accuracy, improve the efficiency of automated testing, and reduce the number of manual interventions.

[0210] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.

[0211] One embodiment of the present invention further provides an image matching device, which is used to implement the above-mentioned embodiments and preferred implementation modes, and will not be repeated hereafter. As used below, the term "unit" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0212] Fig.12 FIG. 1 is a structural block diagram of an image matching device according to one embodiment of the present invention. Fig.12 As shown, the image matching device 120 may include: a first acquiring unit 121 , an extracting unit 122 , a second acquiring unit 123 and a determining unit 124 .

[0213] The first acquisition unit 121 is used to acquire a first image and a second image, wherein the first image is the image to be matched, and the second image is a reference image corresponding to the image to be matched.

[0214] The extraction unit 122 is used to extract the first feature from the first image and the second feature from the second image.

[0215] The second acquisition unit 123 is used to acquire a first target correlation coefficient between the first feature and the second feature, wherein the first target correlation coefficient is used to indicate the degree of correlation between the first feature and the second feature.

[0216] The determining unit 124 is configured to determine, in the second image, a first target region that matches the first image based on the first target correlation coefficient.

[0217] In the image matching device of this embodiment, a first target correlation coefficient is determined by utilizing features of the image to be matched and features of a reference image corresponding to the image to be matched, and then based on the first target correlation coefficient, an area that matches a given image to be matched is found in the reference image corresponding to the image to be matched. This method meets the universality requirements in automated testing scenarios, avoids the need for supervised learning-based image matching methods to pre-manufacture a data set of a certain size and is difficult to produce a universal model, thereby solving the technical problem of low efficiency in image matching and achieving the technical effect of improving the efficiency of image matching.

[0218] It should be noted that the above-mentioned units can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited to this: the above-mentioned units are all located in the same processor; or the above-mentioned units are located in different processors in any combination.

[0219] An embodiment of the present invention further provides a non-volatile storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when executed by a processor.

[0220] Optionally, in this embodiment, the non-volatile storage medium may be configured to store a computer program for performing the following steps:

[0221] S1, obtaining a first image and a second image, wherein the first image is the image to be matched, and the second image is a reference image corresponding to the image to be matched;

[0222] S2, extracting a first feature from the first image, and extracting a second feature from the second image;

[0223] S3, obtaining a first target correlation coefficient between the first feature and the second feature, wherein the first target correlation coefficient is used to indicate the degree of correlation between the first feature and the second feature;

[0224] S4, determining a first target region matching the first image in the second image based on the first target correlation coefficient.

[0225] Optionally, in this embodiment, the above-mentioned non-volatile storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store computer programs.

[0226] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the computer program to execute the steps in any one of the above method embodiments.

[0227] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0228] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0229] S1, obtaining a first image and a second image, wherein the first image is the image to be matched, and the second image is a reference image corresponding to the image to be matched;

[0230] S2, extracting a first feature from the first image, and extracting a second feature from the second image;

[0231] S3, obtaining a first target correlation coefficient between the first feature and the second feature, wherein the first target correlation coefficient is used to indicate the degree of correlation between the first feature and the second feature;

[0232] S4, determining a first target region matching the first image in the second image based on the first target correlation coefficient.

[0233] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0234] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0235] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0236] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0237] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0238] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0239] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.

[0240] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. An image matching method, characterized in that: include: Acquire a first image and a second image, wherein the first image is the image to be matched, and the second image is a reference image corresponding to the image to be matched; extracting a first feature from the first image, and extracting a second feature from the second image; Acquire a first target correlation coefficient between the first feature and the second feature, wherein the first target correlation coefficient is used to indicate a degree of correlation between the first feature and the second feature; determining a first target region matching the first image in the second image based on the first target correlation coefficient; Among them, extracting the first feature from the first image and extracting the second feature from the second image includes: determining that the first image does not include the first text information, or determining that fuzzy matching of the first text information and the second text information in the second image fails, then extracting the first feature from the first image and extracting the second feature from the second image.

2. The method according to claim 1, characterized in that Extracting a first feature from the first image includes: A first target feature map is extracted from the first image based on a feature extraction model, wherein the feature extraction model is obtained based on convolutional neural network training, and the first target feature map includes: shape features of the first image.

3. The method according to claim 2, characterized in that The method further comprises: The first target feature map is output based on a feature output layer of the feature extraction model, wherein the feature output layer is determined based on a size of the first image.

4. The method according to claim 2, characterized in that: Extracting a second feature from the second image includes: A second target feature map is extracted from the second image based on the feature extraction model, wherein the second target feature map includes: shape features of the second image.

5. The method according to claim 4, characterized in that The method further comprises: The second target feature map is output based on a feature output layer of the feature extraction model, wherein the feature output layer is determined based on a size of the first image.

6. The method according to claim 4, characterized in that Obtaining a first target correlation coefficient between the first feature and the second feature includes: Transforming the first target feature map from the time domain to the frequency domain to obtain a third target feature map; Transforming the second target feature map from the time domain to the frequency domain to obtain a fourth target feature map; Normalized cross-correlation processing is performed on the third target feature map and the fourth target feature map to obtain the first target correlation coefficient.

7. The method according to claim 6, characterized in that Performing normalized cross-correlation processing on the third target feature map and the fourth target feature map to obtain the first target correlation coefficient includes: Determine a first complex conjugate value corresponding to the third target characteristic graph and a second complex conjugate value corresponding to the fourth target characteristic graph; Perform inverse Fourier transform on the product of the first complex conjugate value and the second complex conjugate value to obtain the first target correlation coefficient.

8. The method according to claim 4, characterized in that The first target feature map is a multi-channel first target feature map, and the second target feature map is a multi-channel second target feature map. Acquiring a first target correlation coefficient between the first feature and the second feature, comprising: acquiring the first target correlation coefficient between the first target feature graph of each channel and the second target feature graph of each channel, to obtain a plurality of the first target correlation coefficients; Determining a first target region matching the first image in the second image based on the first target correlation coefficient includes: determining the first target region in the second image based on a plurality of the first target correlation coefficients.

9. The method according to claim 8, characterized in that Determining the first target area in the second image based on a plurality of the first target correlation coefficients includes: Obtaining a maximum first target correlation coefficient among a plurality of first target correlation coefficients; adjusting the maximum first target correlation coefficient based on a target adjustment parameter; Determine a correlation coefficient among the plurality of first target correlation coefficients that is greater than or equal to the adjusted maximum first target correlation coefficient as a second target correlation coefficient; The first object region is determined in the second image based on the second object correlation coefficient.

10. The method according to claim 9, characterized in that Determining the first target area in the second image based on the second target correlation coefficient includes: Determine first position information corresponding to the second target correlation coefficient; Converting the first position information to obtain second position information in the second image, wherein the second position information is used to represent position information of a center of the first target area; The first target area is determined based on the second location information.

11. The method according to claim 10, characterized in that Determining second position information in the second image based on the first position information includes: The first position information is converted into the second position information based on the width and height of the first target feature map, the width and height of the second target feature map, and a scaling ratio of the second image to the second target feature map.

12. The method according to claim 10, characterized in that Determining the first target area based on the second location information includes: Determine the second position information as the position information of the center of the first target area; A bounding box of the first target area is determined based on the position information of the center to obtain the first target area.

13. The method according to claim 12, characterized in that In the case where the number of the second target correlation coefficients is multiple, the number of the bounding boxes is multiple, and determining the bounding box of the first target area based on the position information of the center to obtain the first target area includes: Selecting a target bounding box from the plurality of bounding boxes based on an intersection-over-union ratio between the plurality of bounding boxes; Determining that there are multiple target bounding boxes, selecting a first target bounding box from the multiple target bounding boxes based on the target point in the second image; An area within the second image that is framed by the first object boundary is determined as the first object area.

14. The method according to claim 2, characterized in that The receptive field size of the network layer of the feature extraction model does not exceed the size of the second image.

15. The method according to claim 1, characterized in that The method further comprises: Determining that the first image includes the first text information, extracting the first text information from the first image, and extracting the second text information from the second image; Performing fuzzy matching on the first text information and the second text information; If it is determined that the fuzzy matching of the first text information and the second text information is successful, a second target area matching the first text information is determined in the second image.

16. The method according to claim 15, characterized in that Determining a second target area matching the first text information in the second image includes: Determining third position information of the first text information in the second image; The second target area is determined based on the third location information.

17. The method according to any one of claims 1 to 16, characterized in that The method further comprises: Acquire the similarity between the image corresponding to the first target area and the first image; If it is determined that the similarity is greater than a target threshold, a prompt message is output, wherein the prompt message is used to indicate that the first image and the second image are matched successfully.

18. An image matching device, characterized in that: include: A first acquisition unit, configured to acquire a first image and a second image, wherein the first image is an image to be matched, and the second image is a reference image corresponding to the image to be matched; an extraction unit, configured to extract a first feature from the first image and a second feature from the second image; A second acquisition unit, configured to acquire a first target correlation coefficient between the first feature and the second feature, wherein the first target correlation coefficient is used to indicate a correlation degree between the first feature and the second feature; a determining unit, configured to determine, in the second image, a first target region matching the first image based on the first target correlation coefficient; The extraction unit is used to extract the first feature from the first image and the second feature from the second image through the following steps: determining that the first image does not include the first text information, or determining that the fuzzy matching of the first text information and the second text information in the second image fails, then extracting the first feature from the first image and extracting the second feature from the second image.

19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the method described in any one of claims 1 to 17 when executed by a processor.

20. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method described in any one of claims 1 to 17.

Citation Information

Patent Citations

  • Target detection method and device, computer equipment and storage medium

    CN111144398A

  • Image matching method and device, computer equipment and storage medium

    CN111666974A