A Location Recognition Method and Related Devices Based on Picture Object Representation

Through the method based on image object representation, image feature vectors and position coded vectors are obtained, object representation is integrated and optimized, and the problem of low location recognition accuracy caused by viewing angle transformation or object occlusion is solved, and a higher accuracy and robust location recognition is achieved.

CN115393724BActive Publication Date: 2025-07-22CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211153565.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-21
Publication Date
2025-07-22
Estimated Expiration
2042-09-21

AI Technical Summary

Technical Problem

In the prior art, the picture-based location recognition method is inaccurate in the feature description when viewing angle changes or object occlusion, resulting in low recognition accuracy.

Method used

Using a method based on image object representation, the image feature vector and position encoding vector are obtained through the feature extraction module, and after fusing, the object decoding module is input to calculate the confidence of the object representation vector, and the relative position object representation is optimized, and the location is finally determined by matching with the preset reference database.

Benefits of technology

It improves the accuracy and robustness of location recognition, obtains complete feature information of objects in the picture, and makes the feature expression richer and more accurate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115393724B_ABST
    Figure CN115393724B_ABST
Patent Text Reader

Abstract

Embodiments of this application belong to the field of artificial intelligence and relate to a location recognition method and related devices based on picture object representation, including inputting a picture to be recognized into a trained target detection model, obtaining an image feature vector and a position encoding vector through feature extraction, and obtaining a fused feature vector after fusion; obtaining an object representation vector and an object absolute position vector according to the fused feature vector, and calculating the confidence of the object representation vector; obtaining a relative position object representation based on the object representation vector and the object absolute position, using the confidence to optimize the relative position object representation to obtain an optimized object representation, and fusing with the relative position object representation to obtain a complete object feature; matching the complete object feature with the reference object feature of the reference picture in a preset reference database to determine the location of the picture to be recognized. In addition, this application also relates to blockchain technology, and the reference picture can be stored in the blockchain. This application can improve the accuracy of location recognition and has stronger robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a location recognition method and related devices based on picture object representation. Background Art

[0002] In recent years, with the explosive development of data science technology, location recognition based on pictures, as an important branch of its technical route, has received extensive attention in both the academic and industrial fields. Especially in many insurance industries, the retrieval of similar location pictures is of great significance for the judgment of abnormal cases.

[0003] Location recognition, also known as image-based positioning, refers to obtaining a current image and then searching in a pre-constructed environmental map to obtain a most similar reference image, and identifying the geographical location corresponding to the current image according to the geographical location corresponding to the reference image. Currently, there are mainly three branches of picture-based location recognition technology. One is global feature description, one is local feature description, and the other is the combination of global and local feature description. However, the generality of the above methods is poor. In practical applications, the recognition of locations often depends on landmark buildings, distinguishable objects and their relative positions. However, when the perspective changes or the objects are occluded, the above methods lead to inaccurate feature descriptions, thus affecting the accuracy of location recognition. Summary of the Invention

[0004] The purpose of the embodiments of this application is to propose a location recognition method and related devices based on picture object representation to solve the technical problem of low recognition accuracy caused by inaccurate feature description in related technologies.

[0005] To solve the above technical problem, the embodiments of this application provide a location recognition method based on picture object representation, which adopts the following technical solutions:

[0006] Obtain a picture to be recognized, and input the picture to be recognized into a trained target detection model, where the target detection model includes a feature extraction module, an object decoding module, a relative position decoding module and an output module;

[0007] Obtain the image feature vector and position encoding vector of the picture to be recognized through the feature extraction module, and fuse the image feature vector and the position encoding vector to obtain a fused feature vector;

[0008] Input the fused feature vector into the object decoding module to obtain an object representation vector and an object absolute position vector, and calculate the confidence of the object representation vector;

[0009] Input the object representation vector and the absolute position of the object into the relative position decoding module to obtain a relative position object representation, and optimize the relative position object representation according to the confidence level to obtain an optimized object representation;

[0010] Input the relative position object representation and the optimized object representation into the output module for fusion to obtain a complete object feature;

[0011] Match the complete object feature with the reference object feature of the reference picture in the preset reference database to obtain a target reference picture, and determine the location of the picture to be recognized based on the geographical location of the target reference picture.

[0012] Further, the step of obtaining the image feature vector and the position encoding vector of the picture to be recognized by the feature extraction module includes:

[0013] Extract the sub-region features of each sub-region of the picture to be recognized through the feature extraction module, and obtain the image feature vector according to the sub-region features;

[0014] Encode the positions of each sub-region feature according to the positional relationship between the sub-regions to obtain a position encoding vector.

[0015] Further, the object decoding module includes an object embedding layer, an object attention layer, and a decoupled linear layer. The step of inputting the fusion feature vector into the object decoding module to obtain an object representation vector and an object absolute position vector includes:

[0016] Input the trained object encoding into the object embedding layer to generate an object query vector;

[0017] Input the fusion feature vector and the object query vector into the object attention layer to obtain an object global feature;

[0018] Perform decoupling calculation on the object global feature through the decoupled linear layer to obtain an object representation vector and an object absolute position vector.

[0019] Further, the relative position decoding module includes a position embedding layer and a position attention layer. The step of inputting the object representation vector and the object absolute position into the relative position decoding module and calculating to obtain a relative position object representation includes:

[0020] Calculate a relative position matrix according to the object absolute position vector;

[0021] Input the trained position encoding into the position embedding layer to obtain a position query vector;

[0022] Input the position query vector, the object representation vector, and the relative position matrix into the position attention layer for attention calculation to obtain the relative position object representation.

[0023] Further, the step of optimizing the relative position object representation according to the confidence level to obtain the optimized object representation includes:

[0024] Obtain the feature weights of the object representation vector according to the position query vector and the absolute position vector of the object;

[0025] Adjust the feature weights using the confidence level;

[0026] Perform attention calculation on the object representation vector based on the adjusted feature weights to obtain the optimized object representation.

[0027] Further, before the step of inputting the picture to be recognized into the trained object detection model, it further includes:

[0028] Obtain an image data set, and based on the image data set, obtain an image training set and an image validation set. The image data set includes image labels corresponding to each image;

[0029] Input the image training set into a pre-constructed initial object detection model to output a predicted recognition result;

[0030] Iteratively update the initial object detection model based on the predicted recognition result until the model converges to obtain a model to be verified;

[0031] Input the image validation set into the model to be verified for verification to obtain a verification result. When the verification result is greater than or equal to a preset threshold, determine the model to be verified as the object detection model.

[0032] Further, the step of iteratively updating the initial object detection model based on the predicted recognition result until the model converges includes:

[0033] Calculate a loss function based on the predicted recognition result;

[0034] Adjust the model parameters of the initial object detection model based on the loss function, and continue iterative training until the model converges.

[0035] To solve the above technical problems, an embodiment of the present application further provides a location recognition device based on picture object representation, adopting the following technical solutions:

[0036] An acquisition module, configured to acquire a picture to be recognized, and input the picture to be recognized into a trained target detection model, wherein the target detection model includes a feature extraction module, an object decoding module, a relative position decoding module, and an output module;

[0037] The feature extraction module is configured to obtain an image feature vector and a position encoding vector of the picture to be recognized through the feature extraction module, and fuse the image feature vector and the position encoding vector to obtain a fused feature vector;

[0038] The object decoding module is configured to input the fused feature vector into the object decoding module to obtain an object representation vector and an object absolute position vector, and calculate the confidence of the object representation vector;

[0039] The relative position decoding module is configured to input the object representation vector and the object absolute position into the relative position decoding module to obtain a relative position object representation, and optimize the relative position object representation according to the confidence to obtain an optimized object representation;

[0040] The output module is configured to fuse the relative position object representation and the optimized object representation by inputting them into the output module to obtain a complete object feature;

[0041] A matching module is configured to match the complete object feature with a reference object feature of a reference picture in a preset reference database to obtain a target reference picture, and determine the location of the picture to be recognized based on the geographical location of the target reference picture.

[0042] To solve the above technical problems, an embodiment of the present application further provides a computer device, which adopts the following technical solution:

[0043] The computer device includes a memory and a processor. Computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the above-mentioned location recognition method based on picture object representation are implemented.

[0044] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0045] Computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the above-mentioned location recognition method based on picture object representation are implemented.

[0046] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:

[0047] In this application, by obtaining the picture to be recognized and inputting the picture to be recognized into a trained object detection model, where the object detection model includes a feature extraction module, an object decoding module, a relative position decoding module, and an output module; obtaining an image feature vector and a position encoding vector of the picture to be recognized through the feature extraction module, and fusing the image feature vector and the position encoding vector to obtain a fused feature vector; inputting the fused feature vector into the object decoding module to obtain an object representation vector and an object absolute position vector, and calculating the confidence of the object representation vector; inputting the object representation vector and the object absolute position into the relative position decoding module to obtain a relative position object representation, and optimizing the relative position object representation according to the confidence to obtain an optimized object representation; inputting the relative position object representation and the optimized object representation into the output module for fusion to obtain a complete object feature; matching the complete object feature with the reference object feature of the reference picture in the preset reference database to obtain a target reference picture, and determining the location of the picture to be recognized based on the geographical location of the target reference picture; in this application, by decoding and decoupling the fused feature vector obtained by fusing the image feature vector and the position encoding vector, an object representation vector and an object absolute position vector are obtained, then a relative position object representation is obtained according to the object representation vector and the object absolute position vector, and the relative position object representation is optimized using the confidence to obtain an optimized object representation, and the optimized object representation and the relative position object representation are fused to obtain a complete object feature, so that the complete feature information of the object in the picture can be obtained, making the feature expression richer and more accurate, further improving the accuracy of location recognition and having stronger robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the solutions in this application, the following will briefly introduce the drawings required for the description of the embodiments of this application. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0049] Figure 1 is an exemplary system architecture diagram in which this application can be applied;

[0050] Figure 2 is a flowchart of an embodiment of the location recognition method based on picture object representation according to this application;

[0051] Figure 3 is a schematic structural diagram of an embodiment of the location recognition device based on picture object representation according to this application;

[0052] Figure 4 is a schematic structural diagram of an embodiment of the computer device according to this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0054] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0055] To enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0056] This application provides a location recognition method based on picture object representation, which relates to artificial intelligence and can be applied to, for example, Figure 1 the system architecture 100 as shown. The system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0057] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0058] The terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, and so on.

[0059] The server 105 can be a server that provides various services, such as a background server that supports the pages displayed on the terminal devices 101, 102, and 103.

[0060] It should be noted that the location recognition method based on picture object representation provided in the embodiments of the present application is generally executed by the server / terminal device. Correspondingly, the location recognition device based on picture object representation is generally set in the server / terminal device.

[0061] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in

[0062] Continuing to refer to Figure 2 , a flowchart of an embodiment of the location recognition method based on picture object representation according to the present application is shown, including the following steps:

[0063] Step S201, obtain a picture to be recognized, and input the picture to be recognized into a trained target detection model, where the target detection model includes a feature extraction module, an object decoding module, a relative position decoding module, and an output module.

[0064] In this embodiment, the obtained picture to be recognized is input into the trained target detection model, and through the feature extraction module, the object decoding module, the relative position decoding module, and the output module in the target detection model, sequential processing can be performed to output a location recognition result.

[0065] Step S202, obtain an image feature vector and a position encoding vector of the picture to be recognized through the feature extraction module, and fuse the image feature vector and the position encoding vector to obtain a fused feature vector.

[0066] In this embodiment, the image feature extraction module extracts the image features of the picture to be recognized, obtains the image feature vector, and performs position encoding on the image features to obtain the position encoding vector corresponding to the picture to be recognized. Among them, the feature extraction module can use the backbone network extraction method, the convolutional neural network CNN extraction method, or the deep learning network extraction method based on Transformer for feature extraction, which is not limited here.

[0067] Fuse the image feature vector and the position encoding vector to obtain the fused feature vector Vfp, including: flatten the image feature vector, and use the position encoding vector to supplement the features to obtain the fused feature vector. For example, the position encoding vector can be embedded into the image feature vector, the image feature vector and the position encoding vector can be added, or the image feature vector and the position encoding vector can be concatenated to obtain the fused feature vector.

[0068] In this embodiment, the steps of obtaining the image feature vector and the position encoding vector of the picture to be recognized by the above feature extraction module include:

[0069] Extract the sub-region features of each sub-region of the picture to be recognized through the feature extraction module, and obtain the image feature vector according to the sub-region features of each sub-region;

[0070] Encode the positions of each sub-region feature according to the positional relationship between the sub-regions to obtain the position encoding vector.

[0071] Among them, the backbone network extraction method can be used to extract the image features of each sub-region of the picture to be recognized. The backbone network extraction method mainly directly generates a feature map of a specific size by using backbone feature extraction networks such as the Residual Network Resnet and VGG.

[0072] Extract the image features of each sub-region to obtain the sub-region feature Vf corresponding to each sub-region, and concatenate the sub-region features of each sub-region to obtain the image feature vector of the picture to be recognized. This image feature vector contains rich semantic information and accurate position information of the picture to be recognized.

[0073] Encode the positions of the corresponding sub-region features according to the positional relationship between the sub-regions to obtain the position encoding vector. Among them, the position encoding method is fixed position encoding, the dimension of the position encoding vector is the same as the number of sub-region features, and the position encoding vector can be set as a learnable parameter. During the training process of the object detection model, the positional relationship between different sub-region features is obtained through learning.

[0074] In this embodiment, by extracting features from each sub-region of the image to be recognized and performing position encoding according to the positional relationship between sub-regions, the fused feature vector obtained by fusing the image feature vector and the position encoding vector contains both image information and position information, and the image information is more complete.

[0075] Step S203: Input the fused feature vector into the object decoding module to obtain the object representation vector and the object absolute position vector, and calculate the confidence of the object representation vector.

[0076] In this embodiment, during the training of the target detection model, in order to control the computational complexity of the entire model processing process and accelerate the convergence speed, the number of objects extracted by the model is preset to the first preset number On, and the corresponding hyperparameter of the preset number of object decodings is On.

[0077] Input the fused feature vector into the object decoding module to decode it into an object feature vector. The description of the object includes two parts. One is the description of the object itself, and the other is the description of the object position. Then the object feature vector also includes two parts. One is the object representation vector obtained by describing the object itself, and the other is the object absolute position vector obtained by describing the object position. Among them, the object includes, but is not limited to, landmark buildings, objects, plants, etc.

[0078] In this embodiment, the object decoding module includes an object embedding layer, an object attention layer, and a decoupled linear layer. Then, inputting the fused feature vector into the object decoding module for decoding processing includes:

[0079] Input the trained object encoding into the object embedding layer to generate an object query vector;

[0080] Input the fused feature vector and the object query vector into the object attention layer for attention calculation to obtain the object global feature;

[0081] Perform decoupling calculation on the object global feature through the decoupled linear layer to obtain the object representation vector and the object absolute position vector.

[0082] In this embodiment, the object encoding Ol is obtained after training, and the number is On. Input the On object encodings Ol into the object embedding layer to generate On object query vectors Q1 (Query) vectors, and use the fused feature vector Vfp as the V1 (Value) vector and the K1 (Key) vector, and form a QKV matrix vector with the object query vector and input it into the object attention layer for attention calculation to predict the object global feature Of, and the number of the object global feature Of is also On. Among them, the object global feature Of can realize running through the context of the entire image information.

[0083] The attention calculation formula is as follows:

[0084]

[0085] The global features Of of On objects are input into a decoupled linear layer for decoupling calculation to obtain object representation vectors Ovf and object absolute position vectors Opf. Specifically, the decoupled linear layer includes an object representation linear layer and a position representation linear layer. The global features Of of the objects are respectively input into the object representation linear layer and the position representation linear layer to predict On object representation vectors Ovf and On object absolute position vectors Opf.

[0086] It should be understood that during the model training process, the hyperparameter of the preset number of object decodings is On, and the object encoding Ol is obtained through random initialization. As the model is trained, weight features with abstract meanings will be formed. In this embodiment, after the model training is completed, the trained object encoding Ol is obtained as the weight for the attention calculation of the object attention layer of the object decoding module.

[0087] Calculate the confidence Oc of each object representation vector Ovf. A high confidence indicates the existence of an object and a relatively accurate position, while a low confidence indicates that there may be no object or there is a large position deviation even if there is an object.

[0088] The confidence Oc represents the probability of whether there is an object within the predicted bounding box and does not predict which category the object belongs to. The calculation formula of the confidence Oc is as follows:

[0089]

[0090] Among them, Pr(Object) represents the probability of the existence of an object within the predicted bounding box. If there is an object, it is 1; if there is no object, it is 0. represents the IOU (Intersection over Union) between the predicted bounding box and the true bounding box of the object, which reflects the degree of proximity between the predicted bounding box and the true bounding box.

[0091] In this embodiment, after the attention calculation of the fused feature vector, decoupling is performed, and the image features and position features of the picture are processed separately, which can improve the feature processing efficiency, obtain a more accurate feature description, and thus improve the subsequent recognition accuracy.

[0092] Step S204: Input the object representation vector and the object absolute position into the relative position decoding module to obtain the relative position object representation, and optimize the relative position object representation according to the confidence to obtain the optimized object representation.

[0093] In this embodiment, the relationships between objects are constructed through positional relationships. The absolute position vector Opf of an object is a feature matrix with a shape of On*c. The relative position matrix Opfr is obtained through a formula. Oprf is a feature matrix with a shape of On*On*c, and the formula is as follows:

[0094] Opfr i,j = f(Opf i , Opf j )

[0095] Among them, Opf i represents the absolute position vector Opf of the i-th object, and Opf j represents the absolute position vector Opf of the j-th object. Opfr i,j is the relative position matrix of the absolute position vector Opf of the i-th object; the form of f(x) can be selected according to the actual situation. For example, taking the difference, summing, or directly concatenating, etc.

[0096] The relationships between objects are constructed by the relative position decoding module according to the relative position matrix. For object i, the relative position decoding module is used to decode to obtain the relative position object representation.

[0097] In this embodiment, the relative position decoding module includes a position embedding layer and a position attention layer. The trained position encoding is input into the position embedding layer to obtain the position query vector Q2. The position query vector Q2, the object representation vector Ovf, and the relative position matrix Opfr are input into the position attention layer for attention calculation to obtain the relative position object representation.

[0098] Among them, during the training process of the object detection model, the preset hyperparameter of the number of position decodings is Pn, and the hyperparameter Pn is adapted to the hyperparameter On. The relative position decoding module is trained through Pn prior position encodings to obtain the position encoding Pl after training. The position encoding Pl after training can be used as the weight feature for attention calculation in the position attention layer.

[0099] The trained position encoding Pl is input into the position embedding layer for embedding operation to obtain Pn position query vectors Q2. The relative position matrix Opfr i,j is used as the K vector K2, and the object representation vector Ovf i is used as the V vector V2. After attention calculation through the position attention layer, the relative position object representation Orv i of object i is obtained, and the quantity is Pn.

[0100] In step S203, the confidence Oc is calculated for the object representation vector Ovf of each extracted object, and the confidence Oc is used for the relative position object representation Orv iAdjust the feature weights. It should be understood that the feature weights The feature weights adjusted using the confidence Oc

[0101] According to the adjusted weights, perform attention calculation on the object representation vector Ovf of object i i to obtain the optimized object feature Ov i , and the calculation formula is as follows:

[0102]

[0103] In this embodiment, by adjusting the object representation vector using the confidence Oc, the optimized object feature Ov i is obtained, which can make the optimized object feature more reliable and the feature description more accurate.

[0104] Step S205: Input the relative position object representation and the optimized object representation into the output module for fusion to obtain the complete object feature.

[0105] In this embodiment, the relative position object representation Orv i and the optimized object feature Ov i are fused to obtain the enhanced complete object feature. Among them, the fusion operation can be to add the relative position object representation Orv i and the optimized object feature Ov i , or splice the relative position object representation Orv i and the optimized object feature Ov i .

[0106] Step S206: Match the complete object feature with the reference object feature of the reference picture in the preset reference database to obtain the target reference picture, and determine the location of the picture to be recognized based on the geographical location of the target reference picture.

[0107] Location recognition is based on the picture to be recognized obtained, searching in the pre-constructed reference database, and matching the most similar reference picture as the target reference picture. The geographical location corresponding to the target reference picture is the location of the picture to be recognized.

[0108] The target detection model can access the reference database. The reference database includes the reference object features of multiple reference pictures. The reference object feature of each reference picture corresponds to one or more geographical locations. Among them, the reference object features can be obtained in advance and stored in the reference database. After the target detection model outputs the complete object feature, it accesses the reference database and compares the complete object feature with the reference object feature.

[0109] In this embodiment, the complete object features with a confidence Oc greater than or equal to a preset threshold t are selected for matching with the reference database. Specifically, determining the complete object features with a confidence Oc greater than or equal to the preset threshold t includes: determining the object representation vector Ovf with a confidence Oc greater than or equal to the preset threshold t i , as the target object representation vector; determining the target optimized object representation based on the target object representation vector, and then determining the complete target object features; comparing the complete target object features with the reference object features to obtain a comparison result, and determining the location of the picture to be recognized based on the comparison result.

[0110] Among them, comparing the complete target object features with the reference object features to obtain a comparison result, and determining the location of the picture to be recognized based on the comparison result includes: calculating the similarity between the complete target object features and the reference object features, sorting the similarities from largest to smallest, and the geographical location corresponding to the reference object feature with the largest similarity is the location of the picture to be recognized.

[0111] It should be emphasized that to further ensure the privacy and security of the reference pictures, the above reference pictures can also be stored in a node of a blockchain.

[0112] The blockchain referred to in this application is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, in essence, is a decentralized database, a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0113] In this application, by obtaining a fused feature vector from an image feature vector and a position encoding vector, decoding the fused feature vector to obtain an object representation vector and an object absolute position vector, and then decoding the object representation vector and the object absolute position vector to obtain a relative position object representation, and fusing the optimized relative position object representation to obtain the complete object features, the complete feature information of the object in the picture to be recognized can be obtained. The feature expression is richer and more accurate, further improving the accuracy of location recognition and having stronger robustness.

[0114] In some optional implementation manners of this embodiment, before the step of inputting the picture to be recognized into the trained target detection model, it further includes:

[0115] Obtaining an image data set, obtaining an image training set and an image validation set based on the image data set, and the image data set includes an image label corresponding to each image.

[0116] Input the image training set into the pre-built initial target detection model and output the predicted recognition results;

[0117] The initial target detection model is iteratively updated based on the prediction and recognition results until the model converges to obtain a model to be verified;

[0118] The image verification set is input into the model to be verified for verification to obtain a verification result. When the verification result is greater than or equal to a preset threshold, the model to be verified is determined to be a target detection model.

[0119] The image dataset includes multiple images and an image label corresponding to each image, wherein the image label includes the location of an object in the image and the corresponding geographic location.

[0120] After obtaining the image dataset, the image dataset is preprocessed. Data preprocessing includes data cleaning and image enhancement. Data cleaning is to remove invalid images (invalid images can be damaged images, or images with incorrect or missing image labels) and unify the images to the same size. Image enhancement can also be performed on the remaining images after removing invalid images, including random flipping, folding and deformation operations, or adding noise operations, etc., so as to expand the dataset and improve the generalization and accuracy of the model.

[0121] After data preprocessing, the image data set is divided into an image training set and an image verification set according to a preset ratio, for example, image training set: image verification set = 8:2. The image training set is input into the pre-built initial target detection model for training, and the predicted recognition result is output. The process of processing the image training set in the model is shown in steps S202 to S205, which will not be repeated here.

[0122] The loss function is calculated based on the predicted recognition results. The loss function includes position loss and location recognition loss. The weighted sum of position loss and location recognition loss is used to obtain the loss function. Among them, the position loss is the loss obtained based on the ratio of the intersection area and the union area of the predicted bounding box and the true bounding box, that is, the IOU loss; the location recognition loss is the loss of the binary matching arrangement of the true bounding box set and the predicted bounding box set, such as the Hungarian loss. The binary matching arrangement is implemented using the Hungarian algorithm.

[0123] Adjust the model parameters according to the loss function and continue iterative training. When the model is trained to a certain extent, the performance of the model reaches the optimal state and the loss function cannot continue to decrease, that is, convergence occurs. The way to judge convergence is to only calculate the loss function in the previous and next two rounds of iteration. If the loss function is still changing, the image training set can continue to be selected and input into the object detection model for further iterative training. If the loss function does not change significantly, it can be considered that the model converges. At this time, it is determined that the training of the object detection model is completed, and the training is stopped, and the final object detection model is output.

[0124] In this embodiment, adjusting the model parameters based on the loss function can improve the model training speed and ensure the recognition accuracy of the trained model.

[0125] After the model converges, a model to be verified is obtained. The model to be verified is verified by inputting the image verification set into the model to be verified, and the location recognition result is output. Calculate the recognition accuracy according to the location recognition result. If the recognition accuracy is greater than or equal to the preset threshold, the model to be verified is output as the object detection model. If the recognition accuracy is less than the preset threshold, update the training data set and retrain the object detection model.

[0126] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0127] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0128] This application can be used in numerous general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0129] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, read-only memory (ROM), etc., or a random access memory (RAM), etc.

[0130] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit and can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment but can be executed at different moments, and their execution order is not necessarily sequential but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0131] Further reference Figure 3 , as an implementation of the method shown above Figure 2 , this application provides an embodiment of a location recognition device based on picture object representation. This device embodiment corresponds to the method embodiment shown in Figure 2 , and this device can be specifically applied to various electronic devices.

[0132] As Figure 3As shown in the figure, the location recognition device 300 based on picture object representation according to this embodiment includes: an acquisition module 301, a feature extraction module 302, an object decoding module 303, a relative position decoding module 304, an output module 305, and a matching module 306. Among them:

[0133] The acquisition module 301 is used to acquire the picture to be recognized and input the picture to be recognized into a trained target detection model. Among them, the target detection model includes a feature extraction module, an object decoding module, a relative position decoding module, and an output module;

[0134] The feature extraction module 302 is used to obtain the image feature vector and the position encoding vector of the picture to be recognized through the feature extraction module, and fuse the image feature vector and the position encoding vector to obtain a fused feature vector;

[0135] The object decoding module 303 is used to input the fused feature vector into the object decoding module to obtain an object representation vector and an object absolute position vector, and calculate the confidence of the object representation vector;

[0136] The relative position decoding module 304 is used to input the object representation vector and the object absolute position into the relative position decoding module to obtain a relative position object representation, and optimize the relative position object representation according to the confidence to obtain an optimized object representation;

[0137] The output module 305 is used to input the relative position object representation and the optimized object representation into the output module for fusion to obtain a complete object feature;

[0138] The matching module 306 is used to match the complete object feature with the reference object feature of the reference picture in the preset reference database to obtain a target reference picture, and determine the location of the picture to be recognized based on the geographical location of the target reference picture.

[0139] It should be emphasized that to further ensure the privacy and security of the reference picture, the above reference picture can also be stored in a node of a blockchain.

[0140] Based on the above location recognition device based on picture object representation, by decoding and decoupling the fused feature vector obtained by fusing the image feature vector and the position encoding vector, an object representation vector and an object absolute position vector are obtained, and then a relative position object representation is obtained according to the object representation vector and the object absolute position vector, and the relative position object representation is optimized using the confidence to obtain an optimized object representation. By fusing the optimized object representation and the relative position object representation to obtain a complete object feature, the complete feature information of the object in the picture can be obtained, making the feature expression richer and more accurate, further improving the accuracy of location recognition and having stronger robustness.

[0141] In this embodiment, the feature extraction module 302 is further configured to:

[0142] Extract the sub-region features of each sub-region of the picture to be recognized through the feature extraction module, and obtain the image feature vector according to the sub-region features;

[0143] Encode the positions of each sub-region feature according to the positional relationship between the sub-regions to obtain a position encoding vector.

[0144] By extracting features from each sub-region of the picture to be recognized and performing position encoding according to the positional relationship between the sub-regions, the fused feature vector obtained by fusing the image feature vector and the position encoding vector contains both image information and position information, and the image information is more complete.

[0145] In this embodiment, the object decoding module 303 includes an object embedding sub-module, an object attention calculation sub-module, and a decoupling sub-module, where:

[0146] The object embedding sub-module is used to input the trained object encoding into the object embedding layer to generate an object query vector;

[0147] The object attention calculation sub-module is used to input the fused feature vector and the object query vector into the object attention layer for attention calculation to obtain an object global feature;

[0148] The decoupling sub-module is used to perform decoupling calculation on the object global feature through the decoupling linear layer to obtain an object representation vector and an object absolute position vector.

[0149] By performing decoupling after the attention calculation of the fused feature vector, separating the image features and position features of the picture for processing, the feature processing efficiency can be improved, more accurate feature descriptions can be obtained, and thus the subsequent recognition accuracy can be improved.

[0150] In this embodiment, the relative position decoding module 304 includes a calculation sub-module, a position embedding sub-module, and a position attention calculation sub-module, where:

[0151] The calculation sub-module is used to calculate a relative position matrix according to the object absolute position vector;

[0152] The position embedding sub-module is used to input the trained position encoding into the position embedding layer to obtain a position query vector;

[0153] The position attention calculation sub-module is used to input the position query vector, the object representation vector, and the relative position matrix into the position attention layer for attention calculation to obtain a relative position object representation.

[0154] In some alternative implementation manners of this embodiment, the relative position decoding module 304 further includes an optimization sub-module, which is used for:

[0155] Obtaining the feature weight of the object representation vector according to the position query vector and the object absolute position vector;

[0156] Adjusting the feature weight by using the confidence;

[0157] Performing attention calculation on the object representation vector based on the adjusted feature weight to obtain an optimized object representation.

[0158] Adjusting the object representation vector through the confidence Oc to obtain an optimized object feature Ov i , which can make the optimized object feature more reliable and the feature description more accurate.

[0159] In some alternative implementation manners, the above-mentioned location recognition device based on the picture object representation further includes a training module, an update module, and a verification module, where:

[0160] The acquisition module is further used to acquire an image data set, and obtain an image training set and an image verification set based on the image data set, and the image data set includes an image label corresponding to each image;

[0161] The training module is used to input the image training set into a pre-constructed initial target detection model and output a predicted recognition result;

[0162] The update module is used to iteratively update the initial target detection model based on the predicted recognition result until the model converges to obtain a model to be verified;

[0163] The verification module is used to input the image verification set into the model to be verified for verification to obtain a verification result, and when the verification result is greater than or equal to a preset threshold, determine that the model to be verified is the target detection model.

[0164] In this embodiment, by training the target detection model, the target detection process can be simplified and the target detection efficiency can be improved.

[0165] In this embodiment, the update module includes a loss calculation sub-module and an adjustment sub-module, where:

[0166] The loss calculation sub-module is used to calculate a loss function based on the predicted recognition result;

[0167] The adjustment sub-module is used to adjust the model parameters of the initial target detection model based on the loss function and continue iterative training until the model converges.

[0168] Adjusting the model parameters based on the loss function can improve the model training speed and ensure the recognition accuracy of the trained model.

[0169] To solve the above technical problems, the embodiments of the present application also provide a computer device. For details, please refer to Figure 4 , Figure 4 which is the basic structural block diagram of the computer device in this embodiment.

[0170] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that communicate with each other through a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0171] The computer device can be a desktop computer, a notebook, a palm computer, a cloud server and other computing devices. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad or a voice control device, etc.

[0172] The memory 41 includes at least one type of readable storage medium, which includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc. Of course, the memory 41 may also include both the internal storage unit and the external storage device of the computer device 4. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of the location recognition method based on picture object representation. In addition, the memory 41 may also be used to temporarily store various data that have been output or will be output.

[0173] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run the computer-readable instructions stored in the memory 41 or process data, such as running the computer-readable instructions of the location recognition method based on picture object representation.

[0174] The network interface 43 may include a wireless network interface or a wired network interface, and the network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0175] In this embodiment, when the processor executes the computer-readable instructions stored in the memory, the steps of the location recognition method based on picture object representation in the above embodiment are implemented. By decoding and decoupling the fused feature vector obtained by fusing the image feature vector and the position encoding vector, an object representation vector and an object absolute position vector are obtained. Then, a relative position object representation is obtained according to the object representation vector and the object absolute position vector, and the relative position object representation is optimized using confidence to obtain an optimized object representation. The optimized object representation and the relative position object representation are fused to obtain the complete object feature, and the complete feature information of the object in the picture can be obtained, making the feature expression richer and more accurate, further improving the accuracy of location recognition and having stronger robustness.

[0176] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor, so that the at least one processor executes the steps of the location recognition method based on picture object representation as described above. By decoding and decoupling the fused feature vector obtained by fusing the image feature vector and the position encoding vector, an object representation vector and an object absolute position vector are obtained. Then, a relative position object representation is obtained according to the object representation vector and the object absolute position vector, and the relative position object representation is optimized using confidence to obtain an optimized object representation. The optimized object representation and the relative position object representation are fused to obtain the complete object feature, and the complete feature information of the object in the picture can be obtained, making the feature expression richer and more accurate, further improving the accuracy of location recognition and having stronger robustness.

[0177] Through the description of the above embodiments, those skilled in the art can clearly understand that the method of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0178] Obviously, the embodiments described above are only a part of the embodiments of the present application, rather than all the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is similarly within the scope of the patent protection of the present application.

Claims

1. A location recognition method based on picture object representation, characterized in that Including the following steps: Obtain the picture to be recognized, and input the picture to be recognized into a trained object detection model, where the object detection model includes a feature extraction module, an object decoding module, a relative position decoding module, and an output module; Obtain the image feature vector and the position encoding vector of the picture to be recognized through the feature extraction module, and fuse the image feature vector and the position encoding vector to obtain a fused feature vector; Input the fused feature vector into the object decoding module to obtain an object representation vector and an object absolute position vector, and calculate the confidence of the object representation vector; Input the object representation vector and the object absolute position into the relative position decoding module to obtain a relative position object representation, and optimize the relative position object representation according to the confidence to obtain an optimized object representation; Input the relative position object representation and the optimized object representation into the output module for fusion to obtain a complete object feature; Match the complete object feature with the reference object feature of the reference picture in the preset reference database to obtain a target reference picture, and determine the location of the picture to be recognized based on the geographical location of the target reference picture.

2. The method for location recognition based on picture object representation according to claim 1, characterized in that The step of obtaining the image feature vector and the position encoding vector of the picture to be recognized through the feature extraction module includes: Extract the sub-region features of each sub-region of the picture to be recognized through the feature extraction module, and obtain the image feature vector according to the sub-region features; Encode the positions of each sub-region feature according to the positional relationship between the sub-regions to obtain a position encoding vector.

3. The location recognition method based on picture object representation according to claim 1, wherein The object decoding module includes an object embedding layer, an object attention layer, and a decoupled linear layer. The step of inputting the fused feature vector into the object decoding module to obtain an object representation vector and an object absolute position vector includes: Input the trained object encoding into the object embedding layer to generate an object query vector; Input the fused feature vector and the object query vector into the object attention layer for attention calculation to obtain an object global feature; Perform decoupling calculation on the object global feature through the decoupled linear layer to obtain an object representation vector and an object absolute position vector.

4. The method for location recognition based on picture object representation according to claim 1, characterized in that The relative position decoding module includes a position embedding layer and a position attention layer. The step of inputting the object representation vector and the object absolute position into the relative position decoding module and calculating to obtain a relative position object representation includes: Calculate a relative position matrix according to the object absolute position vector; Input the trained position encoding into the position embedding layer to obtain a position query vector; Input the position query vector, the object representation vector, and the relative position matrix into the position attention layer for attention calculation to obtain a relative position object representation.

5. The method for location recognition based on picture object representation according to claim 4, wherein The step of optimizing the relative position object representation according to the confidence to obtain an optimized object representation includes: Obtain the feature weight of the object representation vector according to the position query vector and the object absolute position vector; Adjust the feature weight using the confidence; Performing attention calculation on the object representation vector based on the adjusted feature weights to obtain an optimized object representation.

6. The method for location recognition based on picture object representation according to claim 1, wherein Before the step of inputting the to-be-recognized picture into the trained target detection model, it further includes: Obtaining an image data set, and obtaining an image training set and an image validation set based on the image data set, where the image data set includes an image label corresponding to each image; Inputting the image training set into a pre-constructed initial target detection model and outputting a predicted recognition result; Iteratively updating the initial target detection model based on the predicted recognition result until the model converges to obtain a model to be verified; Inputting the image validation set into the model to be verified for verification to obtain a verification result. When the verification result is greater than or equal to a preset threshold, determining the model to be verified as the target detection model.

7. The method for location recognition based on picture object representation according to claim 6, wherein The step of iteratively updating the initial target detection model based on the predicted recognition result until the model converges includes: Calculating a loss function based on the predicted recognition result; Adjusting the model parameters of the initial target detection model based on the loss function and continuing iterative training until the model converges.

8. A location recognition device based on picture object representation, characterized in that It includes: An acquisition module, configured to acquire a to-be-recognized picture and input the to-be-recognized picture into a trained target detection model, where the target detection model includes a feature extraction module, an object decoding module, a relative position decoding module, and an output module; A feature extraction module, configured to obtain an image feature vector and a position encoding vector of the to-be-recognized picture through the feature extraction module, and fuse the image feature vector and the position encoding vector to obtain a fused feature vector; An object decoding module, configured to input the fused feature vector into the object decoding module to obtain an object representation vector and an object absolute position vector, and calculate the confidence of the object representation vector; A relative position decoding module, configured to input the object representation vector and the object absolute position into the relative position decoding module to obtain a relative position object representation, and optimize the relative position object representation according to the confidence to obtain an optimized object representation; An output module, configured to fuse the relative position object representation and the optimized object representation and input them into the output module to obtain a complete object feature; A matching module, configured to match the complete object feature with a reference object feature of a reference picture in a preset reference database to obtain a target reference picture, and determine the location of the to-be-recognized picture based on the geographical location of the target reference picture.

9. A computer device, including a memory and a processor, where computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of the location recognition method based on picture object representation according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on a computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of the location recognition method based on picture object representation according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Intensive video description method based on position coding fusion

    CN111814844A

  • Method and device for identifying relation between objects and electronic system

    CN112819011A