Method, device and system for determining posture
By detecting text fields in the image to be queried and querying candidate reference images on the server side, the initial position inaccuracy caused by environmental image similarity in underground garages and other places is solved, and a higher position accuracy is achieved.
Patent Information
- Application Number
- CN202010124987.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-08
- Filing Date
- 2020-02-27
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2040-02-27
AI Technical Summary
There are many similar environmental images in the environmental images taken at different locations in underground garages, etc., which makes it difficult for the server to accurately find matching target environmental images, resulting in low accuracy of the initial position.
By detecting and extracting text fields in the image to be queried, and querying candidate reference images on the server based on these text fields, pose solving processing is performed to improve the accuracy of the initial pose.
Even in the presence of more similar image interference, the candidate reference image query based on the text field has high accuracy, and the initial position of the terminal can be determined more accurately.
Smart Images

Figure CN112784174B_ABST
Abstract
Description
[0001] This application claims priority to Chinese patent application No. 201911089900.7, filed on November 8, 2019, entitled “Method, device and system for determining posture”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present disclosure relates to the field of electronic technology, and in particular to a method, device, and system for determining posture. Background Art
[0003] With the development of science and technology, the application of computer vision technology is becoming more and more extensive. Computer vision technology can be used to determine the initial position of the terminal, which includes the position and posture of the terminal in the target location.
[0004] The specific implementation method for determining the initial pose of the terminal is as follows: the terminal captures a query image at its current location in the target location and uploads the query image to a server. The server stores multiple environmental images captured at various locations in the target location, as well as the three-dimensional position information in real space of the physical object points corresponding to each pixel in each environmental image. After receiving the query image sent by the terminal, the server searches the stored environmental images of the target location for a target environmental image that matches the query image. The server also extracts image key points from the query image, identifies target pixels in each pixel in the target environmental image that match the image key points in the query image, and determines the three-dimensional position information in real space of the physical object points corresponding to the image key points in the query image based on the stored three-dimensional position information in real space of the physical object points corresponding to each pixel in each environmental image and the target pixels in the target environmental image that match the image key points in the query image. Finally, the server uses a pose solution algorithm to perform pose solution processing based on the position information of the image key points in the query image and the three-dimensional position information of the corresponding physical object points in real space, thereby obtaining the initial pose of the terminal. The server sends the determined initial position to the terminal, and the terminal performs navigation or other processing based on the initial position.
[0005] Theoretically, in the above process, the closer the shooting location of the target environment image found by the server matches the query image to the shooting location of the query image, the more accurate the final initial pose will be. Of course, there are many factors besides the shooting location that can affect the accuracy of the initial pose, which are not considered here.
[0006] In the process of implementing the present disclosure, the inventors found that there are at least the following problems:
[0007] Because some locations often have similar environments, different environmental images captured at different locations in these locations often contain many similar environmental images. For example, images of underground garages often contain many very similar environmental images. If the target location is an underground garage, when the server searches for a target environmental image that matches the query image, due to the interference of many similar environmental images, it is likely to find an environmental image as the target environmental image that was not actually taken near the location where the query image was captured. However, the initial pose determined based on such a target environmental image has a low accuracy rate. Summary of the Invention
[0008] In order to overcome the problems existing in the related art, the present disclosure provides the following technical solutions:
[0009] According to a first aspect of an embodiment of the present disclosure, a method for determining a posture is provided, the method comprising:
[0010] The terminal obtains an image to be queried at a first position, wherein the scene at the first position includes the scene in the image to be queried; and there is text in the image to be queried;
[0011] Determining N text fields contained in the image to be queried, where N is greater than or equal to 1;
[0012] Sending the N text fields and the image to be queried to a server;
[0013] Receive the initial posture of the terminal at the first position returned by the server; wherein the initial posture is determined by the server according to the N text fields and the image to be queried.
[0014] Through the embodiments of the present disclosure, even if there are many similar environments at different locations in the target location, and there are many similar images in different images taken at different locations in these locations, it is still possible to query candidate reference images that match the image to be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if there is interference from many similar images, the accuracy of the candidate reference images queried based on the text fields is higher, and the initial pose of the terminal obtained by performing pose solution processing based on the candidate reference images with higher accuracy is more accurate.
[0015] In a possible implementation, the terminal obtains the image to be queried at a first location, including:
[0016] capturing a first initial image;
[0017] When no text exists in the first initial image, displaying or announcing a first prompt message; wherein the first prompt message is used to indicate that no text is detected in the first initial image and prompt the user to move the position of the terminal or adjust the shooting angle of the terminal;
[0018] When the terminal captures a second initial image containing text at the first position, the second initial image is determined as the image to be queried.
[0019] Through the above approach, after each user performs an operation, the terminal can evaluate the user's operation based on preset logic and provide appropriate guidance to guide the user to capture a query image with higher image quality. The initial position of the terminal determined based on the query image with higher image quality is more accurate.
[0020] In a possible implementation, the terminal obtains the image to be queried at a first location, including:
[0021] capturing a third initial image;
[0022] Determine a text area image contained in the third initial image by performing text detection processing on the third initial image;
[0023] When the text area image included in the third initial image does not meet the preferred image condition, displaying or announcing a second prompt message; wherein the second prompt message is used to indicate that the text area image included in the third initial image does not meet the preferred image condition and prompting the user to move the terminal in the direction where the physical text is located;
[0024] until the terminal captures a fourth initial image at the first position, the image of the text region containing the image meeting the preferred image condition, and determines the fourth initial image as the image to be queried;
[0025] The preferred image conditions include one or more of the following conditions:
[0026] The size of the text area image is greater than or equal to a size threshold;
[0027] The clarity of the text area image is greater than or equal to a clarity threshold;
[0028] The texture complexity of the text region image is less than or equal to a complexity threshold.
[0029] Through the above approach, after each user performs an operation, the terminal can evaluate the user's operation based on preset logic and provide appropriate guidance to guide the user to capture a query image with higher image quality. The initial position of the terminal determined based on the query image with higher image quality is more accurate.
[0030] In a possible implementation, the terminal obtains the image to be queried at a first location, including:
[0031] capturing a fifth initial image;
[0032] Determining N text fields included in the fifth initial image;
[0033] Acquire M text fields contained in a reference query image, wherein a time interval between a capture time of the reference query image and a capture time of the fifth initial image is less than a time threshold, and M is greater than or equal to 1;
[0034] When any text field included in the fifth initial image is inconsistent with each of the M text fields, a third prompt message is displayed or announced by voice; wherein the third prompt message is used to indicate that an incorrect text field is recognized in the fifth initial image and prompt the user to move the position of the terminal or adjust the shooting angle of the terminal;
[0035] When each text field contained in the sixth initial image photographed by the terminal at the first position belongs to the M text fields, the sixth initial image is determined as the image to be queried.
[0036] Through the above approach, after each user performs an operation, the terminal can evaluate the user's operation based on preset logic and provide appropriate guidance to guide the user to capture a query image with higher image quality. The initial position of the terminal determined based on the query image with higher image quality is more accurate.
[0037] In a possible implementation, the terminal obtains the image to be queried at a first location, including:
[0038] The terminal captures a first image of the current scene at the first position; the first image contains text;
[0039] Performing text detection processing on the first image to obtain at least one text area image;
[0040] The at least one text region image contained in the first image is used as a query image.
[0041] After capturing a first image containing a text region image that meets preset conditions through user guidance, the first image can be cropped or cutout to obtain the text region image within the first image. When subsequently sending the query image to the server, only the cropped or cutout text region image is required, without sending the entire first image.
[0042] In a possible implementation, the method further includes:
[0043] Determine a location area of the text area image in the image to be queried;
[0044] Sending the location area to the server; the initial pose is determined by the server based on the N text fields and the image to be queried, including: the initial pose is determined by the server based on the location area of the text area image in the image to be queried, the N text fields and the image to be queried.
[0045] The terminal sends the complete environment image, and the server uses the text area image to determine the initial pose. Therefore, the terminal can also send the location area of the text area image in the environment image to the server. The server determines the initial pose of the terminal based on the location area.
[0046] In a possible implementation, the method further includes:
[0047] Obtaining location information of the terminal;
[0048] Sending the positioning information to the server; the initial posture is determined by the server based on the N text fields and the image to be queried, including: the initial posture is determined by the server based on the N text fields, the image to be queried and the positioning information.
[0049] In the case where there are multiple candidate reference images, the target reference image can be screened out from the candidate reference images based on the positioning information, and the initial pose determined based on the target reference image has higher accuracy.
[0050] In a possible implementation, after receiving the initial posture returned by the server, the method further includes:
[0051] Obtaining a posture change of the terminal;
[0052] Determine a real-time posture according to the initial posture and the posture change of the terminal.
[0053] Simultaneous localization and mapping (SLAM) tracking technology can save computing overhead. The terminal only needs to send the query image and N text fields to the server once, and the server only needs to return the terminal's initial pose based on the query image and N text fields once. The real-time pose can then be determined based on the initial pose using SLAM tracking technology.
[0054] In a possible implementation, after receiving the initial posture returned by the server, the method further includes:
[0055] Get the preview stream of the current scene;
[0056] Determining, based on the real-time posture, preset media content contained in a digital map corresponding to the scene in the preview stream;
[0057] The media content is rendered in the preview stream.
[0058] If the terminal is a mobile phone or wearable AR device, a virtual scene can be constructed based on the real-time pose. First, the terminal can obtain a preview stream of the current scene. For example, a user can capture a preview stream of the current environment in a shopping mall. Next, the terminal can determine the real-time pose using the method described above. Subsequently, the terminal can obtain a digital map, which records the three-dimensional coordinates of various locations in the world coordinate system. Preset media content exists at these preset three-dimensional coordinates. The terminal can determine the target three-dimensional coordinates corresponding to the real-time pose in the digital map. If the target three-dimensional coordinates contain corresponding preset media content, the preset media content is retrieved. For example, if a user is photographing a target store, the terminal recognizes the real-time pose and determines that the camera is currently facing the target store. The preset media content corresponding to the target store can be retrieved. The preset media content corresponding to the target store can include information about the target store, such as which products are worth purchasing. Based on this, the terminal can render the media content in the preview stream. The user can then view the preset media content corresponding to the target store in a preset area near the image of the target store on the phone. After viewing the preset media content corresponding to the target store, the user will have a general understanding of the target store.
[0059] In a possible implementation, determining N text fields contained in the image to be queried includes:
[0060] Determine all text fields in the query image;
[0061] Input each text field into the pre-trained text classifier to obtain the text type corresponding to each text field;
[0062] Among all text fields, text fields whose text types are preset significant types are determined as N text fields.
[0063] All text fields detected in the query image can be filtered to further extract salient text fields. Salient text fields are identifying fields that clearly or uniquely identify an environment. This execution logic can also be implemented on the server. The initial pose determined based on salient text fields is more accurate.
[0064] According to a second aspect of an embodiment of the present disclosure, a method for determining a posture is provided, the method comprising:
[0065] Receiving an image to be queried and N text fields contained in the image to be queried, sent by a terminal, where N is greater than or equal to 1; the image to be queried is obtained based on an image captured by the terminal at a first position; and the scene at the first position includes the scene in the image to be queried;
[0066] Determining candidate reference images according to the N text fields based on pre-stored correspondences between reference images and text fields;
[0067] Determining an initial posture of the terminal at the first position based on the query image and the candidate reference image;
[0068] The initial posture is sent to the terminal.
[0069] Through the embodiments of the present disclosure, even if there are many similar environments at different locations in the target location, and there are many similar images in different images taken at different locations in these locations, it is still possible to query candidate reference images that match the image to be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if there is interference from many similar images, the accuracy of the candidate reference images queried based on the text fields is higher, and the initial pose of the terminal obtained by performing pose solution processing based on the candidate reference images with higher accuracy is more accurate.
[0070] In a possible implementation, the image to be queried is an environment image captured by a terminal, the reference image is a pre-captured environment image, and the candidate reference image is a candidate environment image in the pre-captured environment image.
[0071] In one possible implementation, the image to be queried is a text area image identified in an environmental image captured by a terminal, the reference image is a text area image identified in an environmental image captured in advance, the candidate reference image is a text area image identified in a candidate environmental image, and the candidate environmental image is a candidate environmental image in the environmental image captured in advance.
[0072] In one possible implementation, when the query image is a text region image recognized in an environment image captured by the terminal, determining the initial position of the terminal at the first position based on the query image and the candidate reference image includes:
[0073] Performing image enhancement processing on the image to be queried to obtain an image to be queried after image enhancement processing;
[0074] Based on the query image after image enhancement processing and the candidate reference image, an initial posture of the terminal at the first position is determined.
[0075] After image enhancement processing is performed on the query image, the accuracy of extracting local image features in the query image can be improved, and thus the accuracy of the initial pose determined based on the local image features in the query image and the candidate reference image is also high.
[0076] In a possible implementation, determining the initial posture of the terminal at the first position based on the query image and the candidate reference image includes:
[0077] Determining a target reference image from the candidate reference images; wherein the scene at the first position includes the scene in the target reference image;
[0078] An initial posture of the terminal at the first position is determined according to the image to be queried and the target reference image.
[0079] Because some candidate reference images are interference images—that is, images not necessarily captured near the location where the query image was first captured, but simply corresponding to text fields that coincide with those in the query image—interference images are also used as candidate reference images to determine the terminal's initial pose, which can affect the accuracy of the initial pose. Therefore, candidate reference images can be screened to determine the target reference image.
[0080] In a possible implementation, determining the initial posture of the terminal at the first position according to the query image and the target reference image includes:
[0081] Determining a 2D-2D correspondence between the query image and the target reference image;
[0082] An initial posture of the terminal at the first position is determined according to the 2D-2D correspondence and a preset 2D-3D correspondence of the target reference image.
[0083] The 2D-2D correspondence may include a 2D-2D correspondence between the environment image captured by the terminal and the pre-captured target environment image, and a 2D-2D correspondence between the text area image recognized in the environment image captured by the terminal and the pre-captured target environment image.
[0084] In a possible implementation, determining a 2D-2D correspondence between the query image and the target reference image includes:
[0085] Determine the image key points in the target reference image that correspond to each image key point in the query image, and obtain a 2D-2D correspondence between the query image and the target reference image.
[0086] In one possible implementation, the 2D-3D image of the target reference image includes three-dimensional position information of physical points corresponding to respective image key points in the target reference image in real space. Determining the initial posture of the terminal at the first position based on the 2D-2D correspondence and a preset 2D-3D correspondence of the target reference image includes:
[0087] The initial posture of the terminal is determined based on the 2D-2D correspondence and the three-dimensional position information of the physical points corresponding to each image key point in the target reference image in the actual space.
[0088] In a possible implementation, the method further includes:
[0089] receiving a position area of the text area image sent by the terminal in the image to be queried;
[0090] The determining of the 2D-2D correspondence between the query image and the target reference image includes:
[0091] Based on the location area, determining a target text area image contained in the image to be queried;
[0092] Acquire a text region image contained in the target reference image;
[0093] A 2D-2D correspondence between the target text region image and the text region image included in the target reference image is determined.
[0094] The 2D-2D correspondence may include a 2D-2D correspondence between a target text region image recognized in an environment image captured by the terminal and a text region image recognized in a target environment image captured in advance.
[0095] In a possible implementation manner, determining a target reference image from the candidate reference images includes:
[0096] Determining the image similarity between each candidate reference image and the query image;
[0097] The candidate reference images whose image similarity is greater than or equal to a preset similarity threshold are determined as target reference images.
[0098] Candidate reference images can be screened based on image similarity.
[0099] In a possible implementation manner, determining a target reference image from the candidate reference images includes:
[0100] Obtaining global image features of each candidate reference image;
[0101] Determining global image features of the image to be queried;
[0102] Determining the distances between the global image features of each candidate reference image and the global image features of the query image;
[0103] The candidate reference image whose distance is less than or equal to the preset distance threshold is determined as the target reference image.
[0104] Candidate reference images can be screened based on global image features.
[0105] In a possible implementation manner, determining a target reference image from the candidate reference images includes:
[0106] receiving positioning information sent by the terminal;
[0107] Obtaining the shooting positions corresponding to each candidate reference image;
[0108] Among the candidate reference images, a target reference image whose shooting position matches the positioning information is determined.
[0109] The positioning information sent by the terminal can be used to assist in screening candidate reference images.
[0110] In a possible implementation, when N is greater than 1, determining a target reference image from the candidate reference images includes:
[0111] Among the candidate reference images, a target reference image including the N text fields is determined.
[0112] In the process of determining candidate reference images corresponding to multiple text fields, each text field can be acquired one by one. Whenever a text field is acquired, the candidate reference image corresponding to the currently acquired text field is determined. In this way, the candidate reference image corresponding to each text field can be determined one by one. To further screen the candidate reference images, a target reference image that includes multiple text fields in the query image can be determined in each candidate reference image. If the same target reference image includes multiple text fields contained in the query image, there is a high probability that the shooting position of the target reference image is similar to the shooting position of the query image, and the accuracy of the initial pose determined based on this target reference image is also high.
[0113] In a possible implementation manner, determining a target reference image from the candidate reference images includes:
[0114] When the number of the candidate reference images is equal to 1, the candidate reference image is determined as the target reference image.
[0115] In a possible implementation, determining the candidate reference images according to the N text fields based on the pre-stored correspondence between the reference images and the text fields includes:
[0116] Inputting N text fields contained in the image to be queried into a pre-trained text classifier to obtain a text type of each text field contained in the image to be queried;
[0117] Determine that the text type is a text field of a preset significant type;
[0118] Based on the pre-stored correspondence between reference images and text fields, a candidate reference image corresponding to the salient type of text field is searched.
[0119] All text fields detected in the query image can be filtered to further extract salient text fields. Salient text fields are identifying fields that clearly or uniquely identify an environment. This execution logic can also be set up on the terminal. The initial pose determined based on salient text fields is more accurate.
[0120] According to a third aspect of an embodiment of the present disclosure, a method for determining a posture is provided, the method comprising:
[0121] The terminal obtains an image to be queried at a first position, wherein the scene at the first position includes the scene in the image to be queried;
[0122] Sending the image to be queried to a server, so that the server determines N text fields contained in the image to be queried, and determining an initial posture of the terminal at the first position according to the N text fields and the image to be queried, where N is greater than or equal to 1;
[0123] Receive the initial position of the terminal at the first position returned by the server.
[0124] In a possible implementation, the method further includes:
[0125] Obtaining location information of the terminal;
[0126] sending the positioning information to the server;
[0127] The determining, according to the N text fields and the image to be queried, an initial posture of the terminal at the first position includes:
[0128] An initial posture of the terminal at the first position is determined according to the N text fields, the image to be queried, and the positioning information.
[0129] In a possible implementation, after receiving the initial posture returned by the server, the method further includes:
[0130] Obtaining a posture change of the terminal;
[0131] Determine a real-time posture according to the initial posture and the posture change of the terminal.
[0132] In a possible implementation, after receiving the initial posture returned by the server, the method further includes:
[0133] Get the preview stream of the current scene;
[0134] Determining, based on the real-time posture, preset media content contained in a digital map corresponding to the scene in the preview stream;
[0135] The media content is rendered in the preview stream.
[0136] According to a fourth aspect of an embodiment of the present disclosure, a method for determining a posture is provided, the method comprising:
[0137] Receiving an image to be queried sent by a terminal, wherein the image to be queried is obtained based on an image captured by the terminal at a first location; and a scene at the first location includes a scene in the image to be queried;
[0138] Determining N text fields contained in the image to be queried, where N is greater than or equal to 1;
[0139] Determining candidate reference images according to the N text fields based on pre-stored correspondences between reference images and text fields;
[0140] Determining an initial posture of the terminal at the first position based on the query image and the candidate reference image;
[0141] The initial posture is sent to the terminal.
[0142] In a possible implementation, determining the initial position of the terminal at the first position based on the query image and the candidate reference image includes:
[0143] Determining a target reference image from the candidate reference images; wherein the scene at the first position includes the scene in the target reference image;
[0144] An initial posture of the terminal at the first position is determined according to the image to be queried and the target reference image.
[0145] In a possible implementation, determining the initial posture of the terminal at the first position according to the query image and the target reference image includes:
[0146] Determining a 2D-2D correspondence between the query image and the target reference image;
[0147] An initial posture of the terminal at the first position is determined according to the 2D-2D correspondence and a preset 2D-3D correspondence of the target reference image.
[0148] In a possible implementation, the method further includes:
[0149] Determine a target text area image contained in the image to be queried;
[0150] Acquire a text region image contained in the target reference image;
[0151] A 2D-2D correspondence between the target text region image and the text region image included in the target reference image is determined.
[0152] In a possible implementation manner, determining a target reference image from the candidate reference images includes:
[0153] Determining the image similarity between each candidate reference image and the query image;
[0154] The candidate reference images whose image similarity is greater than or equal to a preset similarity threshold are determined as target reference images.
[0155] In a possible implementation manner, determining a target reference image from the candidate reference images includes:
[0156] Obtaining global image features of each candidate reference image;
[0157] Determining global image features of the image to be queried;
[0158] Determining the distances between the global image features of each candidate reference image and the global image features of the query image;
[0159] The candidate reference image whose distance is less than or equal to the preset distance threshold is determined as the target reference image.
[0160] In a possible implementation manner, determining a target reference image from the candidate reference images includes:
[0161] receiving positioning information sent by the terminal;
[0162] Obtaining the shooting positions corresponding to each candidate reference image;
[0163] Among the candidate reference images, a target reference image whose shooting position matches the positioning information is determined.
[0164] In a possible implementation, when N is greater than 1, determining a target reference image from the candidate reference images includes:
[0165] Among the candidate reference images, a target reference image including the N text fields is determined.
[0166] In a possible implementation manner, determining a target reference image from the candidate reference images includes:
[0167] When the number of the candidate reference images is equal to 1, the candidate reference image is determined as the target reference image.
[0168] In a possible implementation, determining the candidate reference images according to the N text fields based on the pre-stored correspondence between the reference images and the text fields includes:
[0169] Inputting N text fields contained in the image to be queried into a pre-trained text classifier to obtain a text type of each text field contained in the image to be queried;
[0170] Determine that the text type is a text field of a preset significant type;
[0171] Based on the pre-stored correspondence between reference images and text fields, a candidate reference image corresponding to the salient type of text field is searched.
[0172] According to a fifth aspect of an embodiment of the present disclosure, a device for determining posture is provided, which includes at least one module, and the at least one module is used to implement the method for determining posture provided by the first aspect above.
[0173] According to a sixth aspect of an embodiment of the present disclosure, a device for determining posture is provided, the device comprising at least one module, and the at least one module is used to implement the method for determining posture provided by the second aspect above.
[0174] According to a seventh aspect of an embodiment of the present disclosure, a device for determining posture is provided, the device comprising at least one module, and the at least one module is used to implement the method for determining posture provided by the third aspect above.
[0175] According to an eighth aspect of an embodiment of the present disclosure, a device for determining posture is provided, the device comprising at least one module, and the at least one module is used to implement the method for determining posture provided by the fourth aspect above.
[0176] According to the ninth aspect of an embodiment of the present disclosure, a terminal is provided, which includes a processor, a memory, a transceiver, a camera and a bus, wherein the processor, the memory, the transceiver and the camera are connected via the bus; the camera is used to capture images; the transceiver is used to receive and send data; the memory is used to store computer programs; the processor is used to control the memory, the transceiver and the camera; the processor is configured to execute instructions stored in the memory; the processor implements the method for determining posture provided by the first or third aspect above by executing the instructions.
[0177] According to the tenth aspect of an embodiment of the present disclosure, a server is provided, which includes a processor, a memory, a transceiver and a bus, wherein the processor, the memory and the transceiver are connected via a bus; the transceiver is used to receive and send data; the processor is configured to execute instructions stored in the memory; the processor implements the method for determining posture provided by the second or fourth aspect above by executing the instructions.
[0178] According to an eleventh aspect of an embodiment of the present disclosure, a system for determining posture is provided. The system may include a terminal and a server. The terminal may implement the method described in the first aspect above, and the server may implement the method described in the second aspect above.
[0179] According to the twelfth aspect of the embodiment of the present disclosure, a system for determining posture is provided. The system may include a terminal and a server. The terminal may implement the method described in the third aspect above, and the server may implement the method described in the fourth aspect above.
[0180] According to a thirteenth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, comprising instructions, which, when the computer-readable storage medium is run on a terminal, enables the terminal to execute the method described in the first or third aspect above.
[0181] According to a fourteenth aspect of an embodiment of the present disclosure, a computer program product comprising instructions is provided. When the computer program product is run on a terminal, the terminal executes the method described in the first or third aspect above.
[0182] According to a fifteenth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, comprising instructions, which, when the computer-readable storage medium is run on a server, enables the server to execute the method described in the second or fourth aspect above.
[0183] According to a sixteenth aspect of an embodiment of the present disclosure, a computer program product comprising instructions is provided, which, when run on a server, enables the server to execute the method described in the second or fourth aspect above.
[0184] According to a seventeenth aspect of an embodiment of the present disclosure, a method for determining a posture is provided, the method comprising:
[0185] Obtaining a pre-acquired reference image;
[0186] determining a text field contained in each reference image;
[0187] The text fields and the reference images are stored in correspondence.
[0188] In a possible implementation manner, determining the text field contained in each reference image includes:
[0189] For each reference image, a text detection process is performed on the reference image to determine a text region image contained in the reference image; and a text field contained in the text region image is determined.
[0190] In a possible implementation, the method further includes:
[0191] Determining a 2D-3D correspondence relationship of each text region image based on the 2D points of the text region image contained in each reference image and a pre-acquired 2D-3D correspondence relationship of each reference image;
[0192] The 2D-3D correspondence relationship of each text area image is stored.
[0193] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:
[0194] Through the embodiments of the present disclosure, even if there are many similar environments at different locations in the target location, and there are many similar images in different images taken at different locations in these locations, it is still possible to query candidate reference images that match the image to be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if there is interference from many similar images, the accuracy of the candidate reference images queried based on the text fields is higher, and the initial pose of the terminal obtained by performing pose solution processing based on the candidate reference images with higher accuracy is more accurate.
[0195] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0196] The accompanying drawings, which are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure. In the drawings:
[0197] Figure 1 is a schematic structural diagram of a terminal according to an exemplary embodiment;
[0198] Figure 2 is a schematic structural diagram of an application framework layer according to an exemplary embodiment;
[0199] Figure 3 is a schematic structural diagram of a server according to an exemplary embodiment;
[0200] Figure 4 is a structural diagram of a system for determining a posture according to an exemplary embodiment;
[0201] Figure 5 is a structural diagram of a system for determining a posture according to an exemplary embodiment;
[0202] Figure 6 is a flow chart of a method for determining a posture according to an exemplary embodiment;
[0203] Figure 7 is a flow chart of a method for determining a posture according to an exemplary embodiment;
[0204] Figure 8 is a flow chart of a method for determining a posture according to an exemplary embodiment;
[0205] Figure 9 is a schematic diagram showing a user guidance interface according to an exemplary embodiment;
[0206] Figure 10 is a schematic diagram showing a user guidance interface according to an exemplary embodiment;
[0207] Figure 11 is a schematic diagram showing a user guidance interface according to an exemplary embodiment;
[0208] Figure 12 is a schematic diagram of an underground garage according to an exemplary embodiment;
[0209] Figure 13 is a schematic diagram of a corridor environment according to an exemplary embodiment;
[0210] Figure 14 is a flow chart of a method for determining a posture according to an exemplary embodiment;
[0211] Figure 15 is a structural diagram of a system for determining a posture according to an exemplary embodiment;
[0212] Figure 16 is a flow chart of a method for determining a posture according to an exemplary embodiment;
[0213] Figure 17 is a flow chart of a method for determining a posture according to an exemplary embodiment;
[0214] Figure 18 is a flow chart of a method for determining a posture according to an exemplary embodiment;
[0215] Figure 19 is a flow chart of a method for determining a posture according to an exemplary embodiment;
[0216] Figure 20 is a flow chart of a method for determining a posture according to an exemplary embodiment;
[0217] Figure 21 is a flow chart of a method for determining a posture according to an exemplary embodiment;
[0218] Figure 22 is a flow chart of a method for determining a posture according to an exemplary embodiment;
[0219] Figure 23 is a flowchart illustrating an offline calibration method according to an exemplary embodiment;
[0220] Figure 24 is a flowchart illustrating an offline calibration method according to an exemplary embodiment;
[0221] Figure 25 is a structural schematic diagram of a device for determining a posture according to an exemplary embodiment;
[0222] Figure 26 is a structural schematic diagram of a device for determining a posture according to an exemplary embodiment;
[0223] Figure 27 is a structural diagram of a device for determining a posture according to an exemplary embodiment;
[0224] Figure 28 The figure is a schematic structural diagram of a device for determining a posture according to an exemplary embodiment.
[0225] The above drawings illustrate specific embodiments of the present disclosure, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the present disclosure in any way, but rather to illustrate the concepts of the present disclosure to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0226] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0227] The words "initial", "first", "second", "Nth", "target", "candidate" and so on used in the following embodiments are only used to distinguish different nouns. The above words or similar words do not constitute a limitation on the embodiments of the present disclosure. "First", "second" or similar words are only used to distinguish different nouns and do not constitute a limitation on the order of precedence.
[0228] Figure 1 A schematic structural diagram of the terminal 100 is shown.
[0229] The terminal 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0230] It should be understood that the structures illustrated in the embodiments of the present disclosure do not constitute a specific limitation on the terminal 100. In other embodiments of the present application, the terminal 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0231] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors.
[0232] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.
[0233] Processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 110 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 110 latency, and thus improves system efficiency.
[0234] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.
[0235] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 110 may include multiple I2C bus lines. The processor 110 may be coupled to the touch sensor 180K, the charger, the flash, the camera 193, and the like via different I2C bus interfaces. For example, the processor 110 may be coupled to the touch sensor 180K via the I2C interface, enabling communication between the processor 110 and the touch sensor 180K via the I2C bus interface, thereby implementing the touch function of the terminal 100.
[0236] The I2S interface can be used for audio communication. In some embodiments, the processor 110 can include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface, enabling the function of answering calls through a Bluetooth headset.
[0237] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 170 and the wireless communication module 160 can be coupled via a PCM bus interface. In some embodiments, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, enabling the function of answering calls via a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.
[0238] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface, enabling the function of playing music through Bluetooth headphones.
[0239] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display 194 and the camera 193. MIPI interfaces include the camera serial interface (CSI) and the display serial interface (DSI). In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to implement the camera function of the terminal 100. The processor 110 and the display 194 communicate via the DSI interface to implement the display function of the terminal 100.
[0240] The GPIO interface can be configured via software. The GPIO interface can be configured as either a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 110 to the camera 193, display 194, wireless communication module 160, audio module 170, sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.
[0241] The USB interface 130 is an interface that complies with USB standards and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface 130 can be used to connect a charger to charge the terminal 100 and to transfer data between the terminal 100 and peripheral devices. It can also be used to connect headphones to play audio. This interface can also be used to connect other electronic devices, such as augmented reality devices.
[0242] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present disclosure is merely an illustrative illustration and does not constitute a structural limitation on the terminal 100. In other embodiments of the present application, the terminal 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0243] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the terminal 100. While charging the battery 142, the charging management module 140 can also power the electronic device through the power management module 141.
[0244] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.
[0245] The wireless communication function of the terminal 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.
[0246] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.
[0247] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied on the terminal 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.
[0248] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.
[0249] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. applied on the terminal 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.
[0250] In some embodiments, the antenna 1 of the terminal 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the terminal 100 can communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).
[0251] Terminal 100 implements display functions through a GPU, display screen 194, and an application processor. The GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0252] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, terminal 100 may include one or N display screens 194, where N is a positive integer greater than one.
[0253] The terminal 100 can realize the shooting function through the ISP, camera 193, video codec, GPU, display screen 194 and application processor.
[0254] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.
[0255] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the terminal 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.
[0256] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the terminal 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0257] Video codecs are used to compress or decompress digital video. Terminal 100 may support one or more video codecs. This allows terminal 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0258] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications in the terminal 100, such as image recognition, face recognition, speech recognition, and text comprehension.
[0259] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the terminal 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0260] The internal memory 121 can be used to store computer executable program codes, which include instructions. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the terminal 100 (such as audio data, a phone book, etc.), etc. In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. The processor 110 executes various functional applications and data processing of the terminal 100 by running instructions stored in the internal memory 121 and / or instructions stored in a memory provided in the processor.
[0261] The terminal 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.
[0262] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.
[0263] The speaker 170A, also called a "horn", is used to convert an audio electrical signal into a sound signal. The terminal 100 can listen to music or listen to a hands-free call through the speaker 170A.
[0264] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the terminal 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.
[0265] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The terminal 100 can be provided with at least one microphone 170C. In other embodiments, the terminal 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the terminal 100 can also be provided with three, four or more microphones 170C to realize sound signal collection, noise reduction, and identification of sound sources, and realize directional recording function, etc.
[0266] The headphone jack 170D is used to connect a wired headphone and can be the USB interface 130 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0267] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be located on display screen 194. There are many types of pressure sensors 180A, such as resistive, inductive, and capacitive. A capacitive pressure sensor can include at least two parallel plates made of conductive material. When force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Terminal 100 determines the intensity of the pressure based on this change in capacitance. When a touch operation is applied to display screen 194, terminal 100 detects the touch intensity based on pressure sensor 180A. Terminal 100 can also calculate the touch location based on the detection signal from pressure sensor 180A. In some embodiments, touch operations applied to the same touch location but with different touch intensities can correspond to different operation instructions. For example, when a touch operation with an intensity less than a first pressure threshold is applied to a short message application icon, a command to view short messages is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to a short message application icon, a command to create a new short message is executed.
[0268] The gyroscope sensor 180B can be used to determine the motion posture of the terminal 100. In some embodiments, the angular velocity of the terminal 100 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 180B. The gyroscope sensor 180B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyroscope sensor 180B detects the angle of the terminal 100 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the terminal 100 through reverse movement to achieve anti-shake. The gyroscope sensor 180B can also be used for navigation and somatosensory game scenes.
[0269] The air pressure sensor 180C is used to measure air pressure. In some embodiments, the terminal 100 calculates the altitude using the air pressure value measured by the air pressure sensor 180C to assist in positioning and navigation.
[0270] The magnetic sensor 180D includes a Hall sensor. The terminal 100 can use the magnetic sensor 180D to detect the opening and closing of the flip case. In some embodiments, when the terminal 100 is a flip phone, the terminal 100 can detect the opening and closing of the flip cover based on the magnetic sensor 180D. Furthermore, based on the detected opening and closing status of the case or flip cover, features such as automatic unlocking of the flip cover can be configured.
[0271] Accelerometer 180E can detect the magnitude of acceleration of terminal 100 in all directions (generally three axes). When terminal 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the electronic device's posture, enabling applications such as switching between landscape and portrait modes and pedometers.
[0272] The distance sensor 180F is used to measure distance. The terminal 100 can measure distance using infrared or laser. In some embodiments, when shooting a scene, the terminal 100 can use the distance sensor 180F to measure distance to achieve fast focusing.
[0273] The proximity light sensor 180G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. The terminal 100 emits infrared light outward through the light emitting diode. The terminal 100 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the terminal 100. When insufficient reflected light is detected, the terminal 100 can determine that there is no object near the terminal 100. The terminal 100 can use the proximity light sensor 180G to detect when the user holds the terminal 100 close to the ear to talk, so as to automatically turn off the screen to save power. The proximity light sensor 180G can also be used in leather case mode and pocket mode to automatically unlock and lock the screen.
[0274] Ambient light sensor 180L is used to sense ambient light brightness. Terminal 100 can adaptively adjust the brightness of display screen 194 based on the perceived ambient light brightness. Ambient light sensor 180L can also be used to automatically adjust white balance when taking photos. Ambient light sensor 180L can also work with proximity light sensor 180G to detect whether terminal 100 is in a pocket to prevent accidental touches.
[0275] The fingerprint sensor 180H is used to collect fingerprints. The terminal 100 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, fingerprint photography, fingerprint call answering, etc.
[0276] Temperature sensor 180J is used to detect temperature. In some embodiments, terminal 100 uses the temperature detected by temperature sensor 180J to implement a temperature handling strategy. For example, when the temperature reported by temperature sensor 180J exceeds a threshold, terminal 100 reduces the performance of a processor located near temperature sensor 180J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature falls below another threshold, terminal 100 heats battery 142 to prevent abnormal shutdown of terminal 100 due to low temperature. In other embodiments, when the temperature falls below yet another threshold, terminal 100 boosts the output voltage of battery 142 to prevent abnormal shutdown due to low temperature.
[0277] The touch sensor 180K is also referred to as a "touch-sensitive device." The touch sensor 180K can be disposed on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also referred to as a "touch screen." The touch sensor 180K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations can be provided via the display screen 194. In other embodiments, the touch sensor 180K can also be disposed on the surface of the terminal 100, at a location different from that of the display screen 194.
[0278] The bone conduction sensor 180M can obtain vibration signals. In some embodiments, the bone conduction sensor 180M can obtain vibration signals from the vibrating bones of the human body. The bone conduction sensor 180M can also contact the human pulse to receive blood pressure pulse signals. In some embodiments, the bone conduction sensor 180M can also be set in headphones to form bone conduction headphones. The audio module 170 can parse out voice signals based on the vibration signals of the vibrating bones of the human body obtained by the bone conduction sensor 180M to implement voice functions. The application processor can parse heart rate information based on the blood pressure pulse signals obtained by the bone conduction sensor 180M to implement heart rate detection functions.
[0279] Keys 190 include a power button, a volume button, etc. Keys 190 may be mechanical keys or touch keys. Terminal 100 may receive key inputs and generate key signal inputs related to user settings and function control of terminal 100.
[0280] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, motor 191 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0281] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.
[0282] The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to or removed from the terminal 100 by inserting it into or removing it from the SIM card interface 195. The terminal 100 can support one or N SIM card interfaces, where N is a positive integer greater than one. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The terminal 100 interacts with the network through the SIM card to implement functions such as calls and data communications. In some embodiments, the terminal 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the terminal 100 and cannot be separated from the terminal 100.
[0283] The software system of the terminal 100 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. The embodiment of the present disclosure takes the Android system of the layered architecture as an example to exemplify the software structure of the terminal 100.
[0284] Figure 2 It is a software structure block diagram of the terminal 100 according to an embodiment of the present disclosure.
[0285] A layered architecture divides software into several layers, each with distinct roles and responsibilities. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0286] The application layer can include a series of application packages.
[0287] like Figure 2 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0288] The application framework layer provides an application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0289] like Figure 2 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0290] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.
[0291] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.
[0292] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.
[0293] The phone manager is used to provide communication functions of the terminal 100, such as management of call status (including answering, hanging up, etc.).
[0294] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0295] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically without user interaction. For example, the Notification Manager is used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.
[0296] Android Runtime includes core libraries and a virtual machine. Android runtime is responsible for scheduling and management of the Android system.
[0297] The core library consists of two parts: one is the function that needs to be called by the Java language, and the other is the Android core library.
[0298] The application layer and application framework layer run in a virtual machine. The virtual machine executes Java files in the application layer and application framework layer as binary files. The virtual machine manages object lifecycles, stack management, thread management, security and exception management, and garbage collection.
[0299] The system library can include multiple functional modules, such as surface manager, media library, 3D graphics processing library (such as OpenGL ES), 2D graphics engine (such as SGL), etc.
[0300] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.
[0301] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0302] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0303] A 2D graphics engine is a drawing engine for 2D drawings.
[0304] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.
[0305] The following describes the workflow of the software and hardware of the terminal 100 in conjunction with the capture and photo shooting scene.
[0306] When the touch sensor 180K receives a touch operation, the corresponding hardware interrupt is sent to the kernel layer. The kernel layer processes the touch operation into a raw input event (including touch coordinates, touch operation timestamp, and other information). The raw input event is stored in the kernel layer. The application framework layer obtains the raw input event from the kernel layer and identifies the control corresponding to the input event. For example, if the touch operation is a touch single-click operation and the control corresponding to the single-click operation is the control of the camera application icon, the camera application calls the interface of the application framework layer to start the camera application, and then starts the camera driver by calling the kernel layer to capture a still image or video through the camera 193.
[0307] A method for determining posture provided by an embodiment of the present disclosure can be implemented in combination with the components provided in the above-mentioned terminal 100. For example, communication with the server can be achieved through components such as antenna 1, antenna 2, mobile communication module 150, and wireless communication module 160, such as transmitting the image to be queried, N text fields, and receiving the initial posture returned by the server. Some prompt information to the user can be voice broadcast through the audio module 170, speaker 170A, and headphone interface 170D. Some prompt information to the user can be displayed through the display screen 194. The image to be queried, the environmental image, the initial image, etc. can be photographed by the camera 193. The gyroscope sensor 180B can assist in determining the movement posture of the terminal, etc. The initial posture determination function of the terminal 100 can be implemented by the components provided in the above-mentioned terminal 100 and the method provided in the embodiment of the present disclosure. The above is for example only and is not limiting.
[0308] Yet another exemplary embodiment of the present disclosure provides a server 300 .
[0309] The server 300 may include a processor 310 and a transceiver 320. The transceiver 320 may be connected to the processor 310. Figure 3 As shown. The transceiver 320 may include a receiver and a transmitter, and may be used to receive or send messages or data. The transceiver 320 may be a network card. The server 300 may further include an acceleration component (which may be called an accelerator). When the acceleration component is a network acceleration component, the acceleration component may be a network card. The processor 310 may be the control center of the server 300, connecting various parts of the entire server 300, such as the transceiver 320, using various interfaces and lines. In the present disclosure, the processor 310 may be a central processing unit (CPU). Optionally, the processor 310 may include one or more processing units. The processor 310 may also be a digital signal processor, an application-specific integrated circuit, a field programmable gate array, a GPU or other programmable logic device. The server 300 may further include a memory 330. The memory 330 may be used to store software programs and modules. The processor 310 executes various functional applications and data processing of the server 300 by reading the software codes and modules stored in the memory 330.
[0310] An exemplary embodiment of the present disclosure provides a system for determining a posture, such as Figure 4As shown, the system may include a terminal and a server. The terminal may be a mobile terminal, a human-computer interaction device, or an in-vehicle visual perception device, such as a mobile phone, a sweeper, an intelligent robot, an unmanned vehicle, an intelligent monitor, or an augmented reality (AR) wearable device. Accordingly, the method provided in the embodiments of the present disclosure may be used in application areas such as human-computer interaction, in-vehicle visual perception, augmented reality, intelligent monitoring, unmanned driving, and finding a car in a garage.
[0311] During the movement of the terminal, the image capture component in the terminal can collect the video stream of the target place in real time. The terminal can also extract the image to be queried from the video stream. The image to be queried can be considered as a video frame in the video stream or a text area image extracted from a video frame, etc. The terminal can send the image to be queried to the server, and the server determines the initial posture of the terminal based on the image to be queried. The server can then send the determined initial posture to the terminal. The terminal can determine its position and posture in the target place based on the received initial posture, and perform navigation, route planning, obstacle avoidance and other processing.
[0312] The process of determining the terminal's initial position can be considered an online positioning process. Offline calibration can also be performed before online positioning. This offline calibration process collects environmental images captured at various locations within the target location. Based on these images, a three-dimensional (3D) point cloud of the target location is constructed. The 3D point cloud includes the 3D position information in real space for each physical point corresponding to each pixel in the environmental image.
[0313] like Figure 5 As shown, the modules involved in the online positioning process include image extraction module 501, text box detection module 502, text field recognition module 503, text region image enhancement module 504, feature extraction module 505, image retrieval module 506, text region image feature matching module 507, 2D-3D point matching module 508, and pose estimation module 509. The modules involved in the offline calibration process include offline text box detection module 510, offline text field recognition module 511, text index establishment module 512, offline text region image enhancement module 513, offline feature extraction module 514, 2D-3D point correspondence calibration module 515, and correspondence registration module 516. The modules used in the online positioning process or offline calibration process can be increased or decreased according to actual usage requirements. The functions of the modules mentioned above are different, and the method for determining the pose described later can be implemented in the modules mentioned above.
[0314] The processing flow of the method in the embodiments described below in this disclosure is not limited to the order of the processing steps. The order of the steps can be freely swapped or performed in parallel without violating the laws of nature. The steps between different embodiments can also be freely combined without violating the laws of nature.
[0315] An exemplary embodiment of the present disclosure provides a method for determining a posture, which can be applied to a terminal, such as Figure 6 As shown, the processing flow of the method may include the following steps:
[0316] Step S601: The terminal obtains an image to be queried at a first location.
[0317] Wherein, there is text in the image to be queried. The scene at the first position includes the scene in the image to be queried.
[0318] Step S602: Determine N text fields contained in the image to be queried.
[0319] Wherein, N is greater than or equal to 1.
[0320] Step S603: Send N text fields and the image to be queried to the server, so that the server determines the initial posture of the terminal at the first position according to the N text fields and the image to be queried.
[0321] Step S604: Receive the initial posture returned by the server.
[0322] In one possible implementation, the step of the terminal acquiring the image to be queried at the first position may include: capturing a first initial image; when there is no text in the first initial image, displaying or voice broadcasting a first prompt message; wherein the first prompt message is used to indicate that no text is detected in the first initial image, and prompting the user to move the position of the terminal or adjust the shooting angle of the terminal; until the terminal captures a second initial image with text at the first position, the second initial image is determined as the image to be queried.
[0323] In one possible implementation, the step of the terminal acquiring the image to be queried at the first position may include: capturing a third initial image; determining the text area image contained in the third initial image by performing text detection processing on the third initial image; when the text area image contained in the third initial image does not meet the preferred image condition, displaying or voice broadcasting a second prompt message; wherein the second prompt message is used to indicate that the text area image contained in the third initial image does not meet the preferred image condition, and prompting the user to move the terminal toward the direction where the physical text is located; until the terminal captures a fourth initial image at the first position whose text area image meets the preferred image condition, the fourth initial image is determined as the image to be queried.
[0324] The preferred image conditions include one or more of the following conditions:
[0325] The size of the text area image is greater than or equal to the size threshold;
[0326] The clarity of the text area image is greater than or equal to the clarity threshold;
[0327] The texture complexity of the text region image is less than or equal to the complexity threshold.
[0328] In one possible implementation, the step of the terminal obtaining the image to be queried at the first position may include: shooting a fifth initial image; determining N text fields contained in the fifth initial image; obtaining M text fields contained in the reference query image, wherein the time interval between the acquisition time of the reference query image and the acquisition time of the fifth initial image is less than a duration threshold, and M is greater than or equal to 1; when any text field contained in the fifth initial image is inconsistent with each of the M text fields, displaying or voice broadcasting a third prompt message; wherein the third prompt message is used to indicate that an erroneous text field is recognized in the fifth initial image, and prompting the user to move the position of the terminal or adjust the shooting angle of the terminal; until each text field contained in the sixth initial image shot by the terminal at the first position belongs to the M text fields, the sixth initial image is determined as the image to be queried.
[0329] In one possible implementation, the step of obtaining the image to be queried may include: capturing a first image of the current scene; the first image contains text; performing text detection processing on the first image to obtain at least one text area image; and using the at least one text area image contained in the first image as the image to be queried.
[0330] In a possible implementation, the method provided by the embodiment of the present disclosure may also include: determining the position area of the text area image in the image to be queried; sending the position area to the server; and enabling the server to determine the initial posture of the terminal at the first position based on N text fields and the image to be queried, including: enabling the server to determine the initial posture of the terminal at the first position based on the position area of the text area image in the image to be queried, N text fields and the image to be queried.
[0331] In a possible implementation, the method provided by the embodiment of the present disclosure may also include: obtaining the positioning information of the terminal; sending the positioning information to the server; and enabling the server to determine the initial posture of the terminal based on N text fields and the image to be queried, including: enabling the server to determine the initial posture of the terminal based on the positioning information, N text fields and the image to be queried.
[0332] In a possible implementation, after receiving the initial posture returned by the server, the method provided by the embodiment of the present disclosure may further include: determining a real-time posture according to the initial posture and the posture change of the terminal.
[0333] In one possible implementation, after receiving the initial posture returned by the server, the method provided by the embodiment of the present disclosure may further include: obtaining a preview stream of the current scene; determining, based on the real-time posture, preset media content contained in the digital map corresponding to the scene in the preview stream; and rendering the media content in the preview stream.
[0334] Through the embodiments of the present disclosure, even if there are many similar environments at different locations in the target location, and there are many similar images in different images taken at different locations in these locations, it is still possible to query candidate reference images that match the image to be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if there is interference from many similar images, the accuracy of the candidate reference images queried based on the text fields is higher, and the initial pose of the terminal obtained by performing pose solution processing based on the candidate reference images with higher accuracy is more accurate.
[0335] An exemplary embodiment of the present disclosure provides a method for determining a posture, which can be applied to a server, such as Figure 7 As shown, the processing flow of the method may include the following steps:
[0336] Step S701: receiving an image to be queried and N text fields contained in the image to be queried, sent by a terminal.
[0337] Wherein, N is greater than or equal to 1. The image to be queried is obtained based on an image captured by the terminal at the first position; and the scene at the first position includes the scene in the image to be queried.
[0338] Step S702 : Based on the pre-stored correspondence between reference images and text fields, candidate reference images are determined according to N text fields.
[0339] Step S703: Determine the initial posture of the terminal at the first position based on the image to be queried and the candidate reference images.
[0340] Step S704: Send the determined initial posture to the terminal.
[0341] In one possible implementation, the step of determining the initial posture of the terminal at the first position based on the image to be queried and the candidate reference image may include: determining a target reference image in the candidate reference image; wherein the scene at the first position includes the scene in the target reference image; and determining the initial posture of the terminal at the first position based on the image to be queried and the target reference image.
[0342] In one possible implementation, the step of determining the initial posture of the terminal at the first position based on the image to be queried and the target reference image may include: determining the 2D-2D correspondence between the image to be queried and the target reference image; and determining the initial posture of the terminal at the first position based on the 2D-2D correspondence and the 2D-3D correspondence of the preset target reference image.
[0343] In a possible implementation, the method provided by the embodiment of the present disclosure may also include: receiving the position area of the text area image sent by the terminal in the image to be queried; the step of determining the 2D-2D correspondence between the image to be queried and the target reference image may include: determining the target text area image contained in the image to be queried based on the position area; obtaining the text area image contained in the target reference image; determining the 2D-2D correspondence between the target text area image and the text area image contained in the target reference image.
[0344] In one possible implementation, the step of determining a target reference image among candidate reference images may include: determining the image similarity between each candidate reference image and the query image; and determining a first preset number of target reference images with the highest image similarity among the candidate reference images.
[0345] In one possible implementation, the step of determining a target reference image among candidate reference images may include: obtaining global image features of each candidate reference image; determining global image features of an image to be queried; determining the distances between the global image features of each candidate reference image and the global image features of the image to be queried; and determining, among each candidate reference image, a second preset number of target reference images having the smallest distances.
[0346] In one possible implementation, the step of determining a target reference image among candidate reference images may include: receiving positioning information sent by a terminal; obtaining a shooting position corresponding to each candidate reference image; and determining, among each candidate reference image, a target reference image whose shooting position matches the positioning information.
[0347] In a possible implementation, when N is greater than 1, the step of determining a target reference image from candidate reference images may include: determining a target reference image containing N text fields from each candidate reference image.
[0348] In a possible implementation, the step of determining the target reference image from the candidate reference images may include: when the number of the candidate reference images is equal to 1, determining the candidate reference image as the target reference image.
[0349] In one possible implementation, based on the pre-stored correspondence between reference images and text fields, the step of determining candidate reference images according to N text fields may include: inputting the N text fields contained in the image to be queried into a pre-trained text classifier to obtain the text type of each text field contained in the image to be queried; determining a text field whose text type is a preset significant type; and based on the pre-stored correspondence between reference images and text fields, searching for candidate reference images corresponding to the text field of the significant type.
[0350] Through the embodiments of the present disclosure, even if there are many similar environments at different locations in the target location, and there are many similar images in different images taken at different locations in these locations, it is still possible to query candidate reference images that match the image to be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if there is interference from many similar images, the accuracy of the candidate reference images queried based on the text fields is higher, and the initial pose of the terminal obtained by performing pose solution processing based on the candidate reference images with higher accuracy is more accurate.
[0351] An exemplary embodiment of the present disclosure provides a method for determining a posture, which can be applied to a terminal and implemented by a server. Figure 8 As shown, the processing flow of the method may include the following steps:
[0352] Step S801 : guiding the user to capture an image to be queried at a first position.
[0353] The scene at the first position includes the scene in the image to be queried.
[0354] In implementation, the first position may include any geographical location or spatial location. The scene may refer to the scene or environment in which the terminal device is located when in use, such as a room, or a field. The scene may also be the entire scene or part of the scene that can be captured by the image capture component of the terminal within a preset position range. The scene may also include the environmental background, physical objects in the environment, etc. The specific range and size of the scene are freely defined according to actual needs, and the embodiments of the present disclosure do not limit this. The scene of the first position may refer to the specific scene around the first position, and may include a certain preset geographical range or field of view. The image to be queried may be an image captured by the terminal at the first position, and the scene in the image is consistent with the physical scene. The scene in the image to be queried may be part or all of the scene corresponding to the first position. The first position is not limited to a certain precise position, and the actual position is allowed to have a certain accuracy error.
[0355] If the image to be queried is taken by the user using a terminal, some measures can be taken to ensure the image quality of the image to be queried, and the user can be guided to take the image through the user interface (UI). In the embodiment of the present disclosure, three methods are provided to guide the user to take the image to be queried.
[0356] Optionally, step S801 may include: capturing a first initial image; when there is no text in the first initial image, displaying or voice broadcasting a first prompt message; wherein the first prompt message is used to indicate that no text is detected in the first initial image, and prompting the user to move the position of the terminal or adjust the shooting angle of the terminal; until the terminal captures a second initial image with text at the first position, the second initial image is determined as the image to be queried.
[0357] In practice, when a user arrives at a target location, such as a shopping mall, and if it is the first time for the user to visit the mall and the user wants to view some information about the mall through the mobile phone, the user can take out the mobile phone, stand at the second position, open the camera and aim at a certain environment of the mall to shoot. At this time, the terminal can obtain the first initial image and detect whether there is text in the first initial image. When there is no text in the first initial image, such as Figure 9 As shown, the terminal can pop up a prompt box or directly broadcast the first prompt information by voice, for example, it can display "No text is detected in the current image, please try changing the position or adjusting the shooting angle". After receiving the prompt, the user moves the phone toward the location of the physical text. During the movement, the terminal continuously captures the initial image until the user reaches a suitable first position. The terminal detects the text in the second initial image and can determine the second initial image as the image to be queried and send it to the server. At this time, the terminal can also pop up a prompt box to inform the user that the text has been detected and that information query processing is being performed based on the captured image.
[0358] Optionally, step S801 may include: capturing a third initial image; determining a text area image contained in the third initial image by performing text detection processing on the third initial image; when the text area image contained in the third initial image does not meet the preferred image condition, displaying or voice broadcasting a second prompt message; wherein the second prompt message is used to indicate that the text area image contained in the third initial image does not meet the preferred image condition, and prompting the user to move the terminal in the direction where the physical text is located; until a fourth initial image containing a text area image that meets the preferred image condition is captured at the first position, the fourth initial image is determined as the image to be queried.
[0359] Among them, the preferred image conditions may include one or more of the following conditions: the size of the text area image is greater than or equal to the size threshold, the clarity of the text area image is greater than or equal to the clarity threshold, and the texture complexity of the text area image is less than or equal to the complexity threshold.
[0360] In implementation, the terminal can detect the text area image in the third initial image and determine the size of the text area image. If the size of the text area image is small, it means that the text area image may not be very clear, and thus the current image does not meet the requirements. The terminal can also directly determine the clarity of the text area image. If the clarity is less than the clarity threshold, it means that the text area image may not be very clear, and thus the current image does not meet the requirements. The terminal can also determine the texture complexity of the text area image. If the texture complexity of the text area image is high, it means that there are more texture features in the text area image, which may interfere with the subsequent recognition of the text field in the text area image, and thus the current image does not meet the requirements. Other preferred image conditions can be reasonably set based on the preferred image conditions provided in the embodiment of the present disclosure according to actual needs. When the text area image in the third initial image does not meet one or more of the preferred image conditions, the initial image is retaken.
[0361] In actual application, when a user arrives at a target location such as a shopping mall, if it is the first time for the user to visit the mall and he wants to view some information about the mall through his mobile phone, he can take out his mobile phone, stand at the second position, open the camera and aim at a certain environment of the mall to shoot. At this time, the terminal can obtain the third initial image and detect whether the text area image contained in the third initial image meets the preferred image condition. When the size of the text area image contained in the third initial image is smaller than the size threshold, such as Figure 10 As shown, the terminal can pop up a prompt box or directly broadcast the second prompt information by voice, for example, it can display "The text box detected in the current image is small, please try to move closer to the location of the physical text". After the user receives the prompt, he moves the mobile phone toward the location of the physical text. During the movement, the terminal continuously captures the initial image until the user reaches a suitable first position. The size of the text area image detected by the terminal in the fourth initial image is greater than the size threshold. The fourth initial image can be determined as the image to be queried and sent to the server. At this time, the terminal can also pop up a prompt box to prompt the user that the current image has met the requirements and is performing information query processing based on the captured image. When the clarity and texture complexity of the text area image contained in the third initial image do not meet the requirements, the user can also be prompted based on the above method to guide the user to capture the image to be queried with higher image quality.
[0362] Optionally, step S801 may include: capturing a fifth initial image; determining N text fields contained in the fifth initial image; obtaining M text fields contained in a reference query image, wherein the time interval between the acquisition time of the reference query image and the acquisition time of the fifth initial image is less than a duration threshold, and M is greater than or equal to 1; when any text field contained in the fifth initial image is inconsistent with each of the M text fields, displaying or voice broadcasting a third prompt message; wherein the third prompt message is used to indicate that an erroneous text field is identified in the fifth initial image, and prompting the user to move the terminal or adjust the shooting angle of the terminal; until each text field contained in the sixth initial image captured at the first position belongs to the M text fields, the sixth initial image is determined as the image to be queried.
[0363] In practice, if the fifth initial image contains one text field, a search can be performed to determine whether the text field exists among the M text fields contained in the reference query image. If the text field exists among the M text fields contained in the reference query image, the text field identified in the fifth initial image is the correct text field. If the fifth initial image contains at least two text fields, each text field contained in the fifth initial image can be acquired sequentially. Whenever a text field is acquired, a search is performed to determine whether the currently acquired text field exists among the M text fields contained in the reference query image. If the currently acquired text field exists among the M text fields contained in the reference query image, the currently acquired text field is the correct text field. Since the fifth initial image contains at least two text fields, the above determination needs to be performed multiple times. As long as any text field contained in the fifth initial image does not exist among the M text fields contained in the reference query image, the initial image can be retaken.
[0364] In actual applications, when a user comes to a target place such as a shopping mall, if the user is visiting the mall for the first time and wants to view some information about the mall through a mobile phone, he can take out the mobile phone, stand at the second position, turn on the camera and aim at a certain environment of the mall to shoot. At this time, the terminal can obtain the fifth initial image, and the terminal can detect N text area images in the fifth initial image, and respectively identify the text fields contained in each text area image. Then, the terminal can also judge the accuracy of the text field identified from the fifth initial image based on the upper and lower video frames of the fifth initial image in the video stream. The time interval between the upper and lower video frames and the fifth initial image is less than the duration threshold. For example, taking the example of only one text field being detected in the image, the text fields identified in the previous video and the next video frame of the fifth initial image are both "A35", and the text field identified in the fifth initial image is "A36", which means that the text field identified in the fifth initial image may be an incorrect text field, and then as Figure 11 As shown, a prompt box may pop up or the third prompt information may be directly broadcast by voice, for example, it may be displayed that "the wrong text may have been recognized, please try changing the position or adjusting the shooting angle". After the user receives the prompt, the mobile phone moves toward the location of the physical text. During the movement, the terminal continuously captures the initial image until the user reaches a suitable first position. When each text field recognized by the terminal in the sixth initial image belongs to the text field that can be recognized in the upper and lower video frames, the sixth initial image can be determined as the image to be queried and sent to the server. At this time, the terminal can also pop up a prompt box to prompt the user that the correct text has been detected and that information query processing is being performed based on the captured image.
[0365] Through the above approach, after each user action, the terminal can evaluate the user's action based on pre-set logic and provide appropriate guidance, guiding the user to capture the query image with higher image quality. UI guidance can be used to guide users to capture scenes that are more conducive to text recognition. The final recognition result can be verified by using the text field recognition results of the upper and lower video frames, thereby making text recognition more accurate and improving search accuracy.
[0366] The three methods provided above can be used individually, or in combination with any two or all three. In addition, other methods for guiding users to capture high-quality images to be queried can be added for use together.
[0367] In addition to the above methods, if the query image is captured automatically by the terminal, the terminal's image capture component can capture the target location's video stream in real time while the terminal is moving, and extract the query image from the video stream. The query image can be a specific frame in the video stream. The terminal can also capture the target location's environment image every time a preset capture cycle is reached, using the complete environment image as the query image.
[0368] In the case where the image to be queried is automatically taken by the terminal and is a video frame in the video stream, the initial posture can be determined by sampling frames during the video shooting process. For example, the terminal can shoot 60 video frames per second. In order to reduce the amount of calculation generated in the process of determining the initial posture, the 60 video frames can be sampled at fixed intervals to obtain, for example, 30 video frames, and the initial posture of the terminal is determined based on the 30 video frames. Whenever a video frame is obtained, the method provided in the embodiment of the present disclosure can be executed once to determine the initial posture of the terminal at the moment of capturing the corresponding video frame.
[0369] The terminal captures different images to be queried in different postures, and the initial posture of the terminal can be determined based on the features of the captured image to be queried. Among them, the posture mentioned in this application can refer to the global posture, which includes the current initial position and posture of the terminal. The posture can also be a rotation angle. The position can be expressed as the coordinates of the terminal in the world coordinate system, and the posture can be expressed as the rotation angle of the terminal relative to the world coordinate system. The world coordinate system can be a coordinate system in a preset actual space. The east and north directions in the actual space can be used as the x-axis and y-axis of the coordinate system respectively, and a straight line perpendicular to the horizontal plane surrounded by the x-axis and y-axis and passing through the preset origin is used as the z-axis to establish the world coordinate system.
[0370] Step S802: Determine N text fields contained in the image to be queried.
[0371] Wherein, N is greater than or equal to 1.
[0372] In implementation, after obtaining the image to be queried, N text fields contained in the image to be queried can be identified based on optical character recognition (OCR) technology. Specifically, text detection processing can be performed to determine N text area images in the image to be queried, and the text contained in each text area image can be identified to obtain N text fields. A text field can be one character or multiple characters. For example, a text field can be "A", "3", or "5", and a text field can also be "A35". "A", "3", and "5" represent one character respectively, and "A35" represents three characters. One character or multiple characters can be used as a text field. In a possible implementation method, the characters contained in a continuous area image portion in the image to be queried can be used as a text field. The text field can be stored in the terminal in the form of a character string.
[0373] In the process of identifying the text fields contained in the query image, the terminal can first perform text detection processing on the query image and output the text box position coordinates corresponding to the text area image contained in the query image. In the embodiment of the present disclosure, the text detection processing can be performed based on the deep learning target detection algorithm (Single Shot Multi Box Detector, SSD). Multiple text boxes can be detected in a query image, and each text box corresponds to one text field.
[0374] Assuming the target location is an underground garage, parking area signs can be set on the pillars of the underground garage. Figure 12 As shown in the figure, the current parking area mark "A35" is set on the pillar of the underground garage. When the terminal collects the query image in the current parking area, the collected query image is likely to include the current parking area mark "A35". When the terminal performs text detection on the query image including the current parking area mark "A35", it can output the position coordinates of the text box corresponding to the "A35" area image. Or, as Figure 13 As shown, when the terminal collects the image to be queried in the corridor of a building, the collected image to be queried is likely to include the current floor identification "3B" or "3E". When the terminal performs text detection processing on the image to be queried including the current floor identification "3B" or "3E", it can output the text box position coordinates corresponding to the "3B" or "3E" area image.
[0375] In the embodiment of the present disclosure, Figure 14As shown, after determining the text box position coordinates corresponding to the text region image contained in the query image, the text region image can be extracted from the query image based on the text box position coordinates. Convolutional Neural Network (CNN) feature extraction is performed on the text region image, and the extracted CNN features are then input into a recurrent neural network (Long Short-Term Memory, LSTM) for encoding. The encoded CNN features are then classified, ultimately outputting the text field in the text region image, such as "A35."
[0376] Optionally, the N text fields detected in the query image can be screened to further extract text fields of a significant type. A significant type of text field is an identifying field that can clearly or uniquely identify an environment. The terminal can input the N text fields contained in the query image into a pre-trained text classifier to obtain the text type of each text field contained in the query image, and determine that the text type is a preset significant type of text field. This execution logic can also be set to be completed on the server, that is, the terminal sends all text fields to the server, and the server screens out the significant type of text fields from the N text fields based on similar logic. If the terminal screens out the significant type of text field from the N text fields, then what the terminal ultimately sends to the server includes the query image and the significant type of text field.
[0377] In practice, the target location may contain a large number of text fields, some of which are helpful in identifying the current environment, while some text fields may interfere with the process of identifying the current environment. Text fields that are helpful in identifying the current environment can be used as significant types of text fields. Effective text field collection rules can be formulated in advance. In application, text fields with identification in the target location can be selected as positive samples, such as parking area signs "A32" and "B405" in underground garages. At the same time, text fields that are not identification in the target location can also be selected as negative samples, and the classifier can be trained based on positive and negative samples.
[0378] After extracting N text fields from the query image, each of these N text fields can be fed into a trained classifier. If the classifier outputs a value close to or equal to 1, the current text field is considered salient; if the output value is close to or equal to 0, the current text field is considered non-salient. This helps improve the accuracy of the initial pose determined based on the salient text fields and the query image.
[0379] Step S803: Send N text fields and the image to be queried to the server.
[0380] In implementation, the server can determine the terminal's initial pose based on the N text fields and the query image sent by the terminal, and then return the terminal's initial pose to the terminal. The specific method by which the server determines the terminal's initial pose based on the N text fields and the query image sent by the terminal will be described later.
[0381] Step S804: receiving the initial posture of the terminal at the first position returned by the server.
[0382] In implementation, the terminal can perform navigation, route planning, obstacle avoidance and other processing based on the received initial posture at the first position. The initial posture at the first position is determined by the server in real time based on the image to be queried and the text field sent by the terminal. It should be noted that in the embodiment of the present disclosure, the data sent by the terminal to the server, the data received by the terminal from the server, the data sent by the server to the terminal, and the data received by the server from the terminal can all be carried in the information transmitted between the terminal and the server. The message sent between the server and the terminal is in the form of information, which can carry indication information for indicating some specific content. For example, when the terminal sends N text fields and the image to be queried to the server, the terminal can carry the N text fields and the image to be queried in the indication information and send it to the server.
[0383] Optionally, after receiving the initial posture returned by the server, the method provided by the embodiment of the present disclosure may further include: obtaining the posture change of the terminal; and determining the real-time posture according to the initial posture and the posture change of the terminal.
[0384] In practice, if the initial pose is determined using a query image in a video, the terminal's pose changes can be subsequently determined using simultaneous localization and mapping (SLAM) tracking technology. Based on the initial pose and the terminal's pose changes, the real-time pose is determined. SLAM tracking technology saves computational overhead: the terminal only needs to send the query image and N text fields to the server once, and the server only needs to return the terminal's initial pose based on the query image and N text fields. Based on the initial pose, the real-time pose can be determined using SLAM tracking technology. The real-time pose can be the pose of the terminal at any geographic location within the target location, such as the first, second, or third position. The terminal can perform navigation, route planning, obstacle avoidance, and other processing based on the real-time pose.
[0385] Optionally, in addition to performing navigation, route planning, obstacle avoidance and other processing based on the real-time posture, after receiving the initial posture returned by the server, the method provided by the embodiment of the present disclosure may also include: obtaining a preview stream of the current scene; determining the preset media content contained in the digital map corresponding to the scene in the preview stream based on the real-time posture; and rendering the media content in the preview stream.
[0386] In implementation, if the terminal is a mobile phone or wearable AR device, a virtual scene can be constructed based on the real-time pose. First, the terminal can obtain a preview stream of the current scene. For example, a user can capture a preview stream of the current environment in a shopping mall. Next, the terminal can determine the real-time pose using the method described above. Subsequently, the terminal can obtain a digital map, which records the three-dimensional coordinates of various locations in the world coordinate system. Preset media content exists at these preset three-dimensional coordinates. The terminal can determine the target three-dimensional coordinates corresponding to the real-time pose in the digital map. If the target three-dimensional coordinates contain corresponding preset media content, the preset media content is obtained. For example, if a user is photographing a target store, the terminal recognizes the real-time pose and determines that the camera is currently facing the target store. The preset media content corresponding to the target store can be obtained. The preset media content corresponding to the target store can include information about the target store, such as which products are worth purchasing. Based on this, the terminal can render the media content in the preview stream. At this point, the user can view the preset media content corresponding to the target store in a preset area near the image of the target store on the phone. After the user has viewed the preset media content corresponding to the target store, he or she can have a general understanding of the target store.
[0387] Different digital maps can be set for different places, so that when the user moves to other places, he can also obtain the preset media content corresponding to the real-time posture based on the method of rendering media content provided in the embodiment of the present disclosure, and render the media content in the preview stream.
[0388] The following describes a specific method in which the server determines the initial position of the terminal based on the N text fields sent by the terminal and the image to be queried.
[0389] The server may receive a query image and N text fields contained in the query image sent by a terminal, determine a candidate reference image based on the N text fields based on a pre-stored correspondence between reference images and text fields, determine an initial pose of the terminal at a first position based on the query image and the candidate reference images, and send the determined initial pose to the terminal. Where N is greater than or equal to 1. The query image is obtained based on an image captured by the terminal at the first position, and the scene at the first position includes the scene in the query image.
[0390] In implementation, since the amount of calculation involved in determining the initial posture is large and it also requires a large storage space to execute, the process of actually determining the initial posture of the terminal based on N text fields and the image to be queried can be executed by the server.
[0391] Through the offline calibration process, the server can pre-establish a database, in which reference images taken at various locations in the target location can be stored. Through the offline calibration process, the server can also obtain a pre-established 2D-3D correspondence. The 2D-3D correspondence contains a large number of 3D points and corresponding 2D points in the reference image. Each 3D point corresponds to a physical point in the target location, and each 3D point corresponds to the three-dimensional position information of a corresponding physical point in the actual space. In addition, the server can pre-identify the text fields in each reference image and store the corresponding reference images and text fields. When the image to be queried is a complete environmental image taken by the terminal, the reference image can be a complete environmental image taken in advance at various locations in the target location.
[0392] In order to improve the search speed of the text field, a search index (Global Index) may be established based on the text field in the database, and each reference image, text field, and search index may be stored in correspondence.
[0393] After the server receives the query image and N text fields contained in the query image from the terminal, it can perform a search process on the N text fields contained in the query image based on the search index to determine candidate reference images corresponding to the N text fields. The candidate reference image can be a target environment image from a pre-captured environment image. In the process of determining the candidate reference images corresponding to the N text fields, each text field can be acquired one by one. Whenever a text field is acquired, the candidate reference image corresponding to the currently acquired text field is determined. In this way, the candidate reference image corresponding to each text field can be determined one by one.
[0394] For an image in which two or more text fields in a candidate reference image are identical to the text fields contained in the image to be queried, this type of candidate reference image will be determined twice or multiple times based on the above-mentioned determination method, that is, the first candidate reference image can be determined based on the first text field in the image to be queried, and the first candidate reference image can also be determined based on the second text field in the image to be queried. Therefore, the first candidate reference image includes both the first text field and the second text field. Therefore, a deduplication operation can be performed on all determined candidate reference images to remove candidate reference images that have been determined multiple times. After determining the candidate reference images, the server can determine the initial posture of the terminal based on the image to be queried and the candidate reference images, and send the determined initial posture to the terminal.
[0395] Optionally, the above description is about the case where the terminal sends only one image to be queried and N text fields to the server. If the terminal sends multiple images to be queried and N text fields corresponding to each image to be queried to the server, the initial posture of the terminal can be determined based on each image to be queried and the corresponding N text fields one by one, and finally multiple initial postures can be obtained. Among the multiple initial postures, the target initial posture is determined based on probability statistics, and the target initial posture is sent to the terminal. Among the multiple initial postures, the step of determining the target initial posture based on probability statistics may include: determining the target initial posture that appears the most times among the multiple initial postures.
[0396] Optionally, based on the image to be queried and the candidate reference images, the step of determining the initial posture of the terminal may include: determining a target reference image in the candidate reference images; wherein the scene at the first position includes the scene in the target reference image; and determining the initial posture of the terminal based on the image to be queried and the target reference image.
[0397] In implementation, since some of the candidate reference images are interference images, that is, they are not necessarily images pre-taken near the first position where the image to be queried is taken, but the text field corresponding to the candidate reference image and the text field corresponding to the image to be queried happen to be consistent, the interference image is also used as a candidate reference image to determine the initial posture of the terminal, which will affect the accuracy of the initial posture. Therefore, the candidate reference images can be screened to determine the target reference image. The method provided in the embodiment of the present disclosure introduces four ways of screening candidate reference images, which will be introduced in detail later. When the number of candidate reference images is equal to 1, the candidate reference image can be directly determined as the target reference image. After the target reference image is determined, the initial posture of the terminal can be determined based on the image to be queried and the target reference image.
[0398] Optionally, the step of determining the initial posture of the terminal based on the image to be queried and the target reference image may include: determining the 2D-2D correspondence between the image to be queried and the target reference image; and determining the initial posture of the terminal based on the 2D-2D correspondence and the 2D-3D correspondence of the preset target reference image.
[0399] In implementation, during the offline calibration process, the server can pre-extract local image features of each reference image, and store the corresponding environmental images, local image features, and the aforementioned text fields. Local image features can include image key points, such as corner points and other characteristic pixel points in other images. The server can determine the local image features corresponding to the target reference image based on the correspondence between the reference image and the local image features. After the server receives the image to be queried sent by the terminal, it can extract the local image features of the image to be queried, and perform feature matching on the local image features of the image to be queried and the local image features corresponding to the target reference image, that is, match the image key points corresponding to the image to be queried and the image key points corresponding to the target reference image. The image key points are 2D points, and a 2D-2D correspondence between the image to be queried and the target reference image can be obtained. The 2D-2D correspondence between the image to be queried and the target reference image can include a 2D-2D correspondence between the complete environmental image captured by the terminal and the target environmental image in the pre-captured complete environmental image.
[0400] For example, if there are three image key points in the query image (including A1, B1, and C1), and five image key points in the target reference image (including A2, B2, C2, D2, and E2), feature matching can determine three sets of correspondences: A1-B2, B1-E2, and C1-A2. Of course, the feature matching process in actual applications is much more complex, involving a greater number of image key points. Here, we only use a few image key points as examples. It should be noted that, in theory, the image key points corresponding to the query image and the image key points corresponding to the target environment image should correspond to the same physical points.
[0401] During the offline calibration process, the server can establish a 3D point cloud of the target location based on the reference images taken at various locations in the target location. Each pixel in each reference image corresponds to a 3D point in the 3D point cloud, which can be recorded as an initial 2D-3D correspondence. After the server determines the image key points of each reference image, it can determine the 3D point in the 3D point cloud corresponding to each image key point in the target reference image based on the initial 2D-3D correspondence, which can be recorded as the 2D-3D correspondence of the target reference image. The 2D-3D correspondence of the target reference image can be the 2D-3D correspondence of the target reference image corresponding to the first position of the terminal when capturing the image to be queried. During the online positioning process, the server can determine the 2D-2D correspondence between the query image and the target reference image, that is, the correspondence between the image key points of the query image and the image key points of the target environment image. Then, the server can determine the 3D points corresponding to the image key points of the query image in the 3D point cloud based on the 2D-2D correspondence between the query image and the target reference image, as well as the 2D-3D correspondence of the target reference image.
[0402] After determining the 3D points corresponding to the image key points of the query image in the 3D point cloud, the 3D points corresponding to the image key points of the query image in the 3D point cloud, the position information of each image key point of the query image and the three-dimensional position information of each corresponding 3D point can be input into the posture estimation module. The posture estimation module can solve the posture of the terminal and output the initial posture of the terminal.
[0403] The following describes four methods for screening candidate reference images in the method provided by the embodiments of the present disclosure.
[0404] Optionally, the step of determining the target reference image from the candidate reference images may include: determining the image similarity between each candidate reference image and the query image; and determining the candidate reference image whose image similarity is greater than or equal to a preset similarity threshold as the target reference image.
[0405] During implementation, the server may calculate the image similarity between each candidate reference image and the query image based on a preset image similarity algorithm. Subsequently, among the candidate reference images, images whose image similarity is greater than a preset similarity threshold may be determined as target reference images. The number of target reference images determined may be one or more. Alternatively, the candidate reference images may be sorted in descending order of image similarity, and a preset number of images ranked at the top of the list may be determined as target reference images. The image similarity algorithm may include a K-nearest neighbor algorithm, for example. The preset similarity threshold may be determined based on experience and may be set to a more reasonable value based on experience.
[0406] Optionally, the step of determining a target reference image among candidate reference images may include: obtaining global image features of each candidate reference image; determining global image features of the image to be queried; determining the distance between the global image features of each candidate reference image and the global image features of the image to be queried; and determining a candidate reference image whose distance is less than or equal to a preset distance threshold as a target reference image.
[0407] In implementation, the server can determine the global image features corresponding to each candidate reference image based on the correspondence between the pre-stored reference images and the global image features, and can also extract the global image features corresponding to the image to be queried. The global image features can be data represented in the form of vectors, so the distance between the global image features corresponding to each candidate reference image and the global image features corresponding to the image to be queried can be calculated. The distance can be a distance of the Euclidean distance or other types. After calculating the distance, the candidate reference image whose distance is less than or equal to the preset distance threshold can be determined as the target reference image. The preset distance threshold can be determined based on experience, and can be set to a more reasonable value based on experience. Alternatively, after calculating the distance, the candidate reference images can be sorted in order of distance from small to large, and a preset number of candidate reference images ranked first are selected as the target reference images.
[0408] like Figure 14 As shown, in the method provided by the embodiment of the present disclosure, a VGG (a network structure proposed by the Visual Geometry Group) network can be used to extract global image features. The environment image can be input into the VGG network, and the VGG network can perform CNN feature extraction on the environment image. The VGG network includes multiple network layers, and the output of the second-to-last fully connected layer in the multiple network layers can be selected as the extracted CNN feature. The extracted CNN feature is then normalized by the L2 normalization to obtain a 4096-dimensional normalized feature, which is the global image feature of the environment image. In practical applications, the global image features of the environment image can also be extracted by other methods, which are not limited by the embodiment of the present disclosure.
[0409] like Figure 15As shown, the system provided by the embodiments of the present disclosure may include: a video stream input module 1501, an image extraction module 1502, a text box detection module 1503, a text recognition module 1504, a global feature extraction module 1505, a local feature extraction module 1506, an image retrieval module 1507, a 2D-2D feature matching module 1508, a 2D-3D matching module 1509, and a pose estimation module 1510. The video stream input module 1501 can be used to acquire a video stream, the image extraction module 1502 can be used to extract a video frame from the video stream, the text box detection module 1503 can be used to detect a text region image in the video frame, and the text recognition module 1504 can be used to determine the text field in the text region image. The global feature extraction module 1505 can be used to extract global image features of the video frame, and the local feature extraction module 1506 can be used to extract local image features of the video frame, such as image key points. The operations in the global feature extraction module 1505 and the local feature extraction module 1506 can be performed in parallel. The image retrieval module 1507 can be used to search for a target reference image based on the text field and the global image features of the video frame. The 2D-2D feature matching module 1508 can be used to determine the 2D-2D correspondence between the video frame and the target reference image based on the local image features of the video frame and the local image features of the target reference image. The 2D-3D matching module 1509 can be used to determine the 2D-3D correspondence of the video frame based on the 2D-2D correspondence between the video frame and the target reference image. The pose estimation module 1510 can be used to determine the initial pose based on the 2D-3D correspondence of the video frame.
[0410] The video stream input module 1501, image extraction module 1502, text box detection module 1503, and text recognition module 1504 can be deployed on a terminal in the system. The global feature extraction module 1505, local feature extraction module 1506, image retrieval module 1507, 2D-2D feature matching module 1508, 2D-3D matching module 1509, and pose estimation module 1510 can be deployed on a server in the system. The video stream input module 1501 and image extraction module 1502 can be implemented using the acquisition module 1701 of the terminal-side device, while the text box detection module 1503 and text recognition module 1504 can be implemented using the determination module 1702 of the terminal-side device. The global feature extraction module 1505, local feature extraction module 1506, image retrieval module 1507, 2D-2D feature matching module 1508, 2D-3D matching module 1509, and pose estimation module 1510 can be implemented using the determination module 1802 of the server-side device.
[0411] Optionally, the candidate reference images may be screened based on the terminal's positioning information. The method provided by the disclosed embodiment may further include: the terminal obtaining the terminal's positioning information; and sending the positioning information to the server. Accordingly, the step of determining the target reference image from the candidate reference images in the server may include: receiving the positioning information sent by the terminal; obtaining the shooting positions corresponding to each candidate reference image; and determining, from each candidate reference image, the target reference image whose shooting position matches the positioning information.
[0412] In implementation, the server can divide the target location according to a preset unit area, for example, it can be divided into units of 100m×100m to obtain multiple sub-areas. In the process of dividing the sub-areas, the boundaries of adjacent sub-areas can be overlapped to a certain extent. The server can calibrate the sub-area to which each reference image belongs. During the online positioning process, the terminal can first collect the current positioning information of the terminal based on the Global Positioning System (GPS) or the mobile location service (LBS), and send the positioning information to the server. The server can determine the target sub-area to which the positioning belongs. Then, among each candidate reference image, the target reference image whose shooting position belongs to the same target sub-area can be determined, that is, the target reference image whose shooting position matches the positioning information.
[0413] For example, the third floors of two adjacent buildings both have the sign "301", but there is a certain distance between the two adjacent buildings. Even if the signs are repeated, through positioning, the terminal's position can be first circled within a certain range, and then the target reference image that matches the sign "301" can be searched within the certain range.
[0414] Optionally, when the number of text fields contained in the query image is greater than 1, the step of determining the target reference image from the candidate reference images may include: determining a target reference image containing N text fields from each candidate reference image.
[0415] In practice, in the process of determining candidate reference images corresponding to multiple text fields, each text field can be acquired one by one. Whenever a text field is acquired, the candidate reference image corresponding to the currently acquired text field is determined. In this way, the candidate reference image corresponding to each text field can be determined one by one. In order to further screen the candidate reference images, a target reference image that includes multiple text fields in the query image can be determined in each candidate reference image. If the same target reference image includes multiple text fields contained in the query image, then the probability that the shooting position of the target reference image is close to the shooting position of the query image is very high, and the accuracy of the initial posture determined based on this target reference image is also high.
[0416] Optionally, based on the correspondence between pre-stored reference images and text fields, the step of determining candidate reference images according to N text fields may include: inputting the N text fields contained in the image to be queried into a pre-trained text classifier to obtain the text type of each text field contained in the image to be queried; determining a text field whose text type is a preset significant type; and based on the correspondence between pre-stored reference images and text fields, searching for candidate reference images corresponding to the text field of the significant type.
[0417] In implementation, the process of screening N text fields can be set in a terminal or a server. The target location can contain a large number of text fields, some of which are helpful in helping to identify the current environment, while some text fields will interfere with the process of identifying the current environment. Text fields that are helpful in helping to identify the current environment can be used as significant types of text fields. Effective text field collection rules can be formulated in advance. In application, text fields with identification in the target location can be selected as positive samples. At the same time, text fields that are not identification in the target location can also be selected as negative samples. The classifier is trained based on positive and negative samples.
[0418] The algorithm for recognizing text fields and the segmentation accuracy of templates (such as AI segmentation templates) can be determined based on user needs. This means that several characters can be segmented into a text field. A single character can be considered a text field, or all characters in a continuous image area can be considered a text field, or all characters in an area can be considered a text field.
[0419] After extracting N text fields from the query image, each of these N text fields can be fed into a trained classifier. If the classifier outputs a value close to or equal to 1, the current text field is considered salient; if the output value is close to or equal to 0, the current text field is considered non-salient. This helps improve the accuracy of the initial pose determined based on the salient text fields and the query image.
[0420] Optionally, the step of determining the initial posture of the terminal based on the image to be queried and the target reference image may include: for each target reference image, determining the initial correspondence between the image key points in the image to be queried and the image key points in the target reference image, performing geometric verification processing on each pair of image key points in the initial correspondence, excluding image key points with incorrect matches in the initial correspondence, and obtaining a target correspondence; if the target correspondence contains pairs of image key points greater than or equal to a preset threshold, it indicates that the target reference image and the image to be queried are images collected in the same environment; based on the image to be queried and the target reference image, the initial posture of the terminal is determined.
[0421] In implementation, the initial correspondence between the image key points in the query image and the image key points in the target reference image can be determined, and the initial correspondence includes multiple pairs of image key points. Then, each pair of image key points can be geometrically verified to exclude image key points with incorrect matches in the initial correspondence. For example, the initial correspondence includes a total of 150 pairs of image key points. 30 pairs of image key points can be excluded through geometric verification. These 30 pairs of image key points are not actually matched image key points, and the target correspondence can be obtained. Finally, it can be determined whether the target correspondence contains a number of image key points greater than or equal to a preset threshold. In the embodiment of the present disclosure, taking the preset threshold of 100 as an example, after excluding 30 pairs of image key points from the 150 pairs of image key points, 120 pairs of image key points remain, which are greater than the preset threshold of 100 pairs. Therefore, the target reference image and the query image to which the 120 pairs of remaining image key points belong are images collected in the same environment. If the target correspondence contains image key points with a value less than a preset threshold, it means that the target reference image and the query image are not environmental images captured in the same environment, and thus the target reference image can be omitted from determining the initial pose.
[0422] Through the embodiments of the present disclosure, in some scenes with weak textures or high texture similarity (such as in corridors or where the wall occupies a large area of the image to be queried, etc.), candidate reference images that match the image to be queried can be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if the texture in the image to be queried is weak or there are fewer textures, candidate reference images with higher accuracy can still be queried based on the text fields. The initial posture of the terminal obtained by performing posture solution processing based on the candidate reference images with higher accuracy is more accurate. Retrieval and precise positioning based on text field retrieval and feature matching of text area images can utilize text semantic information in the scene, thereby improving the positioning success rate in some areas with similar textures or repeated textures, and because the 2D-3D correspondence of the reference image of the text area image is utilized, the positioning accuracy is made higher.
[0423] Through the embodiments of the present disclosure, text fields can be imperceptibly integrated into visual features, so the recall rate and precision of image retrieval will be higher, and the process is imperceptible to the user, the positioning process is also more intelligent, and the user experience is better.
[0424] Through the embodiments of the present disclosure, even if there are many similar environments at different locations in the target location, and there are many similar images in different images taken at different locations in these locations, it is still possible to query candidate reference images that match the image to be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if there is interference from many similar images, the accuracy of the candidate reference images queried based on the text fields is higher, and the initial pose of the terminal obtained by performing pose solution processing based on the candidate reference images with higher accuracy is more accurate.
[0425] Based on the same inventive concept as the above disclosed embodiment, an exemplary embodiment of the present disclosure provides a method for determining a posture, which can be applied to a terminal and implemented by a server, such as Figure 16 As shown, the processing flow of the method may include the following steps:
[0426] Step S1601 : guiding the user to capture a first image containing a text area image that meets a preset condition at a first position.
[0427] The scene at the first position includes the scene in the image to be queried.
[0428] In implementation, the first position may include any geographical location or spatial location. The scene may refer to the scene or environment in which the terminal device is located when in use, such as a room, or a field. The scene may also be the entire scene or part of the scene that can be captured by the image capture component of the terminal within a preset position range. The scene may also include the environmental background, physical objects in the environment, etc. The specific range and size of the scene are freely defined according to actual needs, and the embodiments of the present disclosure do not limit this. The scene of the first position may refer to the specific scene around the first position, and may include a certain preset geographical range or field of view. The image to be queried may be an image captured by the terminal at the first position, and the scene in the image is consistent with the physical scene. The scene in the image to be queried may be part or all of the scene corresponding to the first position. The first position is not limited to a certain precise position, and the actual position is allowed to have a certain accuracy error.
[0429] If the first image is captured by a user using a terminal, measures can be taken to ensure the image quality of the text area image contained in the captured first image. User guidance can be provided through the user interface. The first image can be an image of the entire environment captured by the terminal, and the text area image can be extracted from the first image. In the disclosed embodiments, three methods are provided for guiding users in obtaining high-quality text area images.
[0430] Optionally, step S1601 may include: capturing a first initial image; when there is no text in the first initial image, displaying or voice broadcasting a first prompt message; wherein the first prompt message is used to indicate that no text is detected in the first initial image, and prompting the user to move the position of the terminal or adjust the shooting angle of the terminal; until the terminal captures a second initial image with text at the first position, the text area image contained in the second initial image is determined as the image to be queried.
[0431] In practice, when a user arrives at a target location, such as a shopping mall, and if it is the first time for the user to visit the mall and the user wants to view some information about the mall through the mobile phone, the user can take out the mobile phone, stand at the second position, open the camera and aim at a certain environment of the mall to shoot. At this time, the terminal can obtain the first initial image and detect whether there is text in the first initial image. When there is no text in the first initial image, such as Figure 9As shown, the terminal can pop up a prompt box or directly broadcast the first prompt information by voice, for example, it can display "No text is detected in the current image, please try changing the position or adjusting the shooting angle". After the user receives the prompt, the mobile phone moves towards the location of the physical text. During the movement, the terminal continuously captures the initial image until the user reaches a suitable first position. The terminal detects the text in the second initial image and can determine the text area image contained in the second initial image as the image to be queried and send it to the server. At this time, the terminal can also pop up a prompt box to prompt the user that the text has been detected and that information query processing is being performed based on the captured image.
[0432] If there is no text in the first initial image, it means that there is no text area image in the first initial image, so the first initial image does not meet the requirements.
[0433] Optionally, step S1601 may include: capturing a third initial image; determining a text area image contained in the third initial image by performing text detection processing on the third initial image; when the text area image contained in the third initial image does not meet the preferred image condition, displaying or voice broadcasting a second prompt message; wherein the second prompt message is used to indicate that the text area image contained in the third initial image does not meet the preferred image condition, and prompting the user to move the terminal in the direction where the physical text is located; until a fourth initial image containing a text area image that meets the preferred image condition is captured at the first position, the text area image contained in the fourth initial image is determined as the image to be queried.
[0434] Among them, the preferred image conditions may include one or more of the following conditions: the size of the text area image is greater than or equal to the size threshold, the clarity of the text area image is greater than or equal to the clarity threshold, and the texture complexity of the text area image is less than or equal to the complexity threshold.
[0435] In implementation, the terminal can detect the text area image in the third initial image and determine the size of the text area image. If the size of the text area image is small, it means that the text area image may not be very clear, and thus the current image does not meet the requirements. The terminal can also directly determine the clarity of the text area image. If the clarity is less than the clarity threshold, it means that the text area image may not be very clear, and thus the current image does not meet the requirements. The terminal can also determine the texture complexity of the text area image. If the texture complexity of the text area image is high, it means that there are more texture features in the text area image, which may interfere with the subsequent recognition of the text field in the text area image, and thus the current image does not meet the requirements. Other preferred image conditions can be reasonably set based on the preferred image conditions provided in the embodiment of the present disclosure according to actual needs. When the text area image in the third initial image does not meet one or more of the preferred image conditions, the initial image is retaken.
[0436] In actual application, when a user arrives at a target location such as a shopping mall, if it is the first time for the user to visit the mall and he wants to view some information about the mall through his mobile phone, he can take out his mobile phone, stand at the second position, open the camera and aim at a certain environment of the mall to shoot. At this time, the terminal can obtain the third initial image and detect whether the text area image contained in the third initial image meets the preferred image condition. When the size of the text area image contained in the third initial image is smaller than the size threshold, such as Figure 10 As shown, the terminal can pop up a prompt box or directly broadcast the second prompt information by voice, for example, it can display "The text box detected in the current image is small, please try to move closer to the location of the physical text". After the user receives the prompt, he moves the mobile phone toward the location of the physical text. During the movement, the terminal continuously captures the initial image until the user reaches a suitable first position. The size of the text area image detected by the terminal in the fourth initial image is greater than the size threshold. The text area image contained in the fourth initial image can be determined as the image to be queried and sent to the server. At this time, the terminal can also pop up a prompt box to prompt the user that the current image meets the requirements and that information query processing is being performed based on the captured image. When the clarity and texture complexity of the text area image contained in the third initial image do not meet the requirements, the user can also be prompted based on the above method to guide the user to capture an image to be queried with higher image quality.
[0437] Optionally, step S1601 may include: capturing a fifth initial image; determining N text fields contained in the fifth initial image; obtaining M text fields contained in a reference query image, wherein the time interval between the acquisition time of the reference query image and the acquisition time of the fifth initial image is less than a duration threshold, and M is greater than or equal to 1; when any text field contained in the fifth initial image is inconsistent with each of the M text fields, displaying or voice broadcasting a third prompt message; wherein the third prompt message is used to indicate that an erroneous text field is identified in the fifth initial image, and prompting the user to move the terminal or adjust the shooting angle of the terminal; until each text field contained in the sixth initial image captured at the first position belongs to the M text fields, the text area image contained in the sixth initial image is determined as the image to be queried.
[0438] In practice, if the fifth initial image contains one text field, a search can be performed to determine whether the text field exists among the M text fields contained in the reference query image. If the text field exists among the M text fields contained in the reference query image, the text field identified in the fifth initial image is the correct text field. If the fifth initial image contains at least two text fields, each text field contained in the fifth initial image can be acquired sequentially. Whenever a text field is acquired, a search is performed to determine whether the currently acquired text field exists among the M text fields contained in the reference query image. If the currently acquired text field exists among the M text fields contained in the reference query image, the currently acquired text field is the correct text field. Since the fifth initial image contains at least two text fields, the above determination needs to be performed multiple times. As long as any text field contained in the fifth initial image does not exist among the M text fields contained in the reference query image, the initial image can be retaken.
[0439] In actual applications, when a user comes to a target place such as a shopping mall, if the user is visiting the mall for the first time and wants to view some information about the mall through a mobile phone, he can take out the mobile phone, stand at the second position, turn on the camera and aim at a certain environment of the mall to shoot. At this time, the terminal can obtain the fifth initial image, and the terminal can detect N text area images in the fifth initial image, and respectively identify the text fields contained in each text area image. Then, the terminal can also judge the accuracy of the text field identified from the fifth initial image based on the upper and lower video frames of the fifth initial image in the video stream. The time interval between the upper and lower video frames and the fifth initial image is less than the duration threshold. For example, taking the example of only one text field being detected in the image, the text fields identified in the previous video and the next video frame of the fifth initial image are both "A35", and the text field identified in the fifth initial image is "A36", which means that the text field identified in the fifth initial image may be an incorrect text field, and then as Figure 11 As shown, a prompt box may pop up or the third prompt information may be directly broadcast by voice, for example, it may be displayed that "the wrong text may have been recognized, please try changing the position or adjusting the shooting angle". After the user receives the prompt, the mobile phone moves toward the location of the physical text. During the movement, the terminal continuously captures the initial image until the user reaches a suitable first position. When each text field recognized by the terminal in the sixth initial image belongs to the text field that can be recognized in the upper and lower video frames, the text area image contained in the sixth initial image can be determined as the image to be queried and sent to the server. At this time, the terminal can also pop up a prompt box to prompt the user that the correct text has been detected and that information query processing is being performed based on the captured image.
[0440] When the sixth initial image includes N text area images, each text area image corresponds to a text field. When each text field contained in the sixth initial image is a text field that can be recognized in the upper and lower video frames, all N text area images can be determined as query images and sent to the server. Alternatively, when the target text field contained in the sixth initial image is a text field that can be recognized in the upper and lower video frames, the text area images corresponding to the target text fields can be determined as query images and sent to the server. If the sixth initial image contains erroneous text fields that are not identifiable in the upper and lower video frames, the text area images corresponding to these erroneous text fields may not be sent to the server.
[0441] Through the above methods, after each user performs an operation, the terminal can evaluate the user's operation based on preset logic and provide appropriate guidance to guide the user to capture the query image with high image quality. The three methods provided above can be used individually, in combination with any two or all three. In addition, other methods for guiding the user to capture high-quality query images can also be used together.
[0442] In addition to the above methods, if the query image is automatically captured by the terminal, the terminal's image capture component can capture a video stream of the target location in real time while the terminal is moving, extract a frame from the video stream, and identify the text area image in the frame as the query image. The query image can be a text area image in a specific frame in the video stream. The terminal can also capture an environmental image of the target location every time a preset capture cycle is reached, and use the text area image in the environmental image as the query image.
[0443] The terminal obtains different query images at different postures, and the initial posture of the terminal can be determined based on the features of the obtained query images, where the initial posture includes the current initial position and posture of the terminal.
[0444] Step S1602: Perform text detection processing on the first image to obtain at least one text region image, and determine the at least one text region image contained in the first image as a query image.
[0445] In implementation, after capturing a first image containing a text region image that meets preset conditions through user guidance, the first image can be cropped or cutout to obtain the text region image within the first image. When subsequently sending the query image to the server, only the cropped or cutout text region image is required, without sending the entire first image.
[0446] Step S1603: Determine N text fields contained in the image to be queried.
[0447] Wherein, N is greater than or equal to 1.
[0448] In implementation, the image to be queried can be a text area image recognized in the environment image taken by the terminal, and there can be multiple text area images, so the image to be queried can be multiple text area images, each text area image corresponds to a text field, so the image to be queried can contain multiple text fields. Of course, if the image to be queried is a text area image, the image to be queried can also contain only one text field.
[0449] A text field can contain one character or multiple characters. For example, a text field can be "A," "3," "5," or even "A35." "A," "3," and "5" each represent one character, and "A35" represents three characters. One character or multiple characters can constitute a text field. In one possible implementation, a text area image is a continuous image, and the characters it contains can constitute a text field. Text fields can be stored in the terminal as character strings.
[0450] Optionally, the N text fields detected in the query image can be screened to further extract text fields of a significant type. A significant type of text field is an identifying field that can clearly or uniquely identify an environment. The terminal can input the N text fields contained in the query image into a pre-trained text classifier to obtain the text type of each text field contained in the query image, and determine that the text type is a preset significant type of text field. This execution logic can also be set to be completed on the server, that is, the terminal sends all text fields to the server, and the server screens out the significant type of text fields from the N text fields based on similar logic. If the terminal screens out the significant type of text field from the N text fields, then what the terminal finally sends to the server includes the significant type of text field and the corresponding text area image.
[0451] The text fields based on the salient types and the corresponding text region images are more conducive to improving the accuracy of the determined initial pose.
[0452] Step S1604: Send N text fields and the image to be queried to the server.
[0453] In implementation, the server may determine the initial posture of the terminal at the first position based on the N text fields and the image to be queried sent by the terminal, and then return the initial posture of the terminal to the terminal.
[0454] Step S1605: receiving the initial posture of the terminal at the first position returned by the server.
[0455] In implementation, the terminal can perform navigation, route planning, obstacle avoidance, etc. based on the received initial position at the first position. The initial position at the first position is determined by the server in real time based on the query image and text field sent by the terminal.
[0456] Optionally, after receiving the initial posture returned by the server, the method provided by the embodiment of the present disclosure may further include obtaining the posture change of the terminal; and determining the real-time posture based on the initial posture and the posture change of the terminal.
[0457] In practice, if the initial pose is determined by a query image in the video, the pose change of the terminal can be determined through the simultaneous localization and mapping (SLAM) tracking technology. Based on the initial pose and the pose change of the terminal, the real-time pose is determined.
[0458] Optionally, in addition to performing navigation, route planning, obstacle avoidance and other processing based on the real-time posture, after receiving the initial posture returned by the server, the method provided by the embodiment of the present disclosure may also include: obtaining a preview stream of the current scene; determining the preset media content contained in the digital map corresponding to the scene in the preview stream based on the real-time posture; and rendering the media content in the preview stream.
[0459] In implementation, if the terminal is a mobile phone or wearable AR device, a virtual scene can be constructed based on the real-time pose. First, the terminal can obtain a preview stream of the current scene. For example, a user can capture a preview stream of the current environment in a shopping mall. Next, the terminal can determine the real-time pose using the method described above. Subsequently, the terminal can obtain a digital map, which records the three-dimensional coordinates of various locations in the world coordinate system. Preset media content exists at these preset three-dimensional coordinates. The terminal can determine the target three-dimensional coordinates corresponding to the real-time pose in the digital map. If the target three-dimensional coordinates contain corresponding preset media content, the preset media content is obtained. For example, if a user is photographing a target store, the terminal recognizes the real-time pose and determines that the camera is currently facing the target store. The preset media content corresponding to the target store can be obtained. The preset media content corresponding to the target store can include information about the target store, such as which products are worth purchasing. Based on this, the terminal can render the media content in the preview stream. At this point, the user can view the preset media content corresponding to the target store in a preset area near the image of the target store on the phone. After the user has viewed the preset media content corresponding to the target store, he or she can have a general understanding of the target store.
[0460] Different digital maps can be set for different places, so that when the user moves to other places, he can also obtain the preset media content corresponding to the real-time posture based on the method of rendering media content provided in the embodiment of the present disclosure, and render the media content in the preview stream.
[0461] The following describes a specific method in which the server determines the initial position of the terminal based on the N text fields sent by the terminal and the image to be queried.
[0462] The server may receive a query image and N text fields contained in the query image sent by a terminal, determine a candidate reference image based on the N text fields based on a pre-stored correspondence between reference images and text fields, determine an initial pose of the terminal based on the query image and the candidate reference images, and send the determined initial pose to the terminal. Where N is greater than or equal to 1. The query image is obtained based on an image captured by the terminal at a first position, and the scene at the first position includes the scene in the query image.
[0463] In practice, through the process of offline calibration, the server can pre-establish a database in which environmental images pre-taken at various locations in the target location can be stored. Through the process of offline calibration, the server can also pre-determine the text area image in each pre-taken environmental image and obtain a pre-established 2D-3D correspondence. The 2D-3D correspondence includes a large number of 3D points and corresponding 2D points in the environmental image. Each 3D point corresponds to a physical point near the location of the physical text in the target location, and each 3D point corresponds to the three-dimensional position information of a corresponding physical point in the actual space. In addition, the server can pre-identify the text fields in the text area images in each pre-taken environmental image and store the corresponding text area images and text fields. When the image to be queried is the text area image in the complete environmental image taken by the terminal, the reference image can be the text area image identified in the pre-taken environmental image.
[0464] In order to improve the search speed of the text field, a search index (Global Index) can be established based on the text field in the database, and the pre-identified text area images, text fields and search indexes are stored in correspondence.
[0465] After the server receives the query image and N text fields contained in the query image sent by the terminal, it can perform retrieval processing on the N text fields contained in the query image based on the search index to determine candidate reference images corresponding to the N text fields. The candidate reference images can be images of text regions identified in the candidate environment image, which can be images in a pre-captured environment image. In the process of determining the candidate reference images corresponding to the N text fields, each text field can be acquired one by one. Whenever a text field is acquired, the candidate reference image corresponding to the currently acquired text field is determined. In this way, the candidate reference image corresponding to each text field can be determined one by one.
[0466] Optionally, based on the image to be queried and the candidate reference images, the step of determining the initial posture of the terminal at the first position may include: determining a target reference image in the candidate reference images; wherein the scene at the first position includes the scene in the target reference image; and determining the initial posture of the terminal at the first position based on the image to be queried and the target reference image.
[0467] In implementation, since some of the candidate reference images are interference images, that is, they are not necessarily images pre-taken near the first position where the image to be queried is taken, but the text field corresponding to the candidate reference image and the text field corresponding to the image to be queried happen to be consistent, the interference image is also used as a candidate reference image to determine the initial posture of the terminal, which will affect the accuracy of the initial posture. Therefore, the candidate reference images can be screened to determine the target reference image. The method provided in the embodiment of the present disclosure introduces three ways of screening candidate reference images, which will be introduced in detail later. When the number of candidate reference images is equal to 1, the candidate reference image can be directly determined as the target reference image. After determining the target reference image, the initial posture of the terminal can be determined based on the image to be queried and the target reference image.
[0468] Optionally, the step of determining the initial posture of the terminal at the first position based on the image to be queried and the target reference image may include: determining the 2D-2D correspondence between the image to be queried and the target reference image; and determining the initial posture of the terminal at the first position based on the 2D-2D correspondence and the 2D-3D correspondence of the preset target reference image.
[0469] In implementation, the 2D-2D correspondence between the query image and the target reference image may include a 2D-2D correspondence between a text region image in a complete environment image captured by the terminal and a target text region image identified in a target environment image in a pre-captured environment image. The 2D-3D correspondence of the target reference image may include a 2D-3D correspondence of the target text region image.
[0470] After determining the 2D-2D correspondence between the query image and the target reference image, the 2D-2D correspondence and the 2D-3D correspondence of the target reference image can be input into the pose estimation module. The pose estimation module can solve the initial pose of the terminal and output the initial pose of the terminal.
[0471] Optionally, before determining the 2D-2D correspondence between the query image and the target reference image, the query image may be subjected to image enhancement processing to obtain the query image after image enhancement processing, and then the 2D-2D correspondence between the query image after image enhancement processing and the target reference image may be determined.
[0472] In implementations, image enhancement processing is performed on the query image to make the determined 2D-2D correspondence between the query image and the target reference image more accurate. The 2D-2D correspondence between the query image and the target reference image may include a correspondence between image key points of the query image and the target reference image. After image enhancement processing is performed on the query image, the image key points of the query image are more accurately determined, and thus the determined 2D-2D correspondence between the query image and the target reference image is more accurate.
[0473] The following describes three methods for screening candidate reference images in the method provided by the embodiments of the present disclosure.
[0474] Optionally, the step of determining the target reference image from the candidate reference images may include: determining the image similarity between each candidate reference image and the query image; and determining the candidate reference image whose image similarity is greater than or equal to a preset similarity threshold as the target reference image.
[0475] In implementation, the server can calculate the image similarity between the candidate text area images identified in the candidate environmental images in each pre-captured environmental image and the text area image in the environmental image captured by the terminal based on a preset image similarity algorithm, and then determine, among the candidate text area images, the images whose image similarity is greater than a preset similarity threshold as target text area images. The number of target text area images determined can be one or more. Alternatively, the candidate text area images can be sorted in descending order of image similarity, and a preset number of images with the highest sorting scores can be determined as target text area images. The image similarity algorithm can include a K-nearest neighbor algorithm, etc. The preset similarity threshold can be determined based on experience, and can be set to a more reasonable value based on experience.
[0476] Optionally, the step of determining a target reference image among candidate reference images may include: obtaining global image features of each candidate reference image; determining global image features of the image to be queried; determining the distance between the global image features of each candidate reference image and the global image features of the image to be queried; and determining a candidate reference image whose distance is less than or equal to a preset distance threshold as a target reference image.
[0477] In implementation, the server can determine the global image features corresponding to the candidate text area images identified in each candidate environmental image based on the correspondence between the text area images identified in the pre-stored environmental image and the global image features, and can also extract the global image features corresponding to the environmental image captured by the terminal. The global image features can be data represented in the form of vectors, so the distance between the global image features corresponding to each candidate text area image and the global image features corresponding to the environmental image captured by the terminal can be calculated. The distance can be a distance of the type such as Euclidean distance. After calculating the distance, the candidate text area images whose distance is less than or equal to the preset distance threshold can be determined as the target text area image. The preset distance threshold can be determined based on experience and can be set to a more reasonable value based on experience. Alternatively, after calculating the distance, the candidate text area images can be sorted in order of distance from small to large, and a preset number of candidate text area images ranked first are selected as the target text area images.
[0478] Optionally, the candidate reference images may be screened based on the terminal's positioning information. The method provided by the disclosed embodiment may further include: the terminal obtaining the terminal's positioning information; and sending the positioning information to the server. Accordingly, the step of determining the target reference image from the candidate reference images in the server may include: receiving the positioning information sent by the terminal; obtaining the shooting positions corresponding to each candidate reference image; and determining, from each candidate reference image, the target reference image whose shooting position matches the positioning information.
[0479] Optionally, based on the correspondence between pre-stored reference images and text fields, the step of determining candidate reference images according to N text fields may include: inputting the N text fields contained in the image to be queried into a pre-trained text classifier to obtain the text type of each text field contained in the image to be queried; determining a text field whose text type is a preset significant type; and based on the correspondence between pre-stored reference images and text fields, searching for candidate reference images corresponding to the text field of the significant type.
[0480] In practice, the process of screening N text fields can be performed on a terminal or a server. A target location may contain a large number of text fields, some of which are helpful in identifying the current environment, while others interfere with the process. Text fields that are helpful in identifying the current environment can be considered salient text fields. Based on salient text fields and the query image, the accuracy of the initial pose determined can be improved.
[0481] Optionally, the step of determining the initial posture of the terminal based on the image to be queried and the target reference image may include: for each target reference image, determining the initial correspondence between the image key points in the image to be queried and the image key points in the target reference image, performing geometric verification processing on each pair of image key points in the initial correspondence, excluding image key points with incorrect matches in the initial correspondence, and obtaining a target correspondence; if the target correspondence contains pairs of image key points greater than or equal to a preset threshold, it indicates that the target reference image and the image to be queried are images collected in the same environment; based on the image to be queried and the target reference image, the initial posture of the terminal is determined.
[0482] In implementation, an initial correspondence between image keypoints in a query image and image keypoints in a target reference image can be determined, where the initial correspondence includes multiple pairs of image keypoints. A geometric verification process can then be performed on each pair of image keypoints to eliminate any incorrectly matched image keypoints in the initial correspondence.
[0483] The processing of some terminals and the processing of servers in the embodiments disclosed herein are similar to the processing of some terminals and the processing of servers in the previously disclosed embodiments. Some of the shared features are not described in detail in the embodiments disclosed herein, and reference may be made to the description of the processing of terminals and the processing of servers in the previously disclosed embodiments.
[0484] Through the embodiments of the present disclosure, in some scenes with weak textures or high texture similarity (such as in corridors or where the wall occupies a large area of the image to be queried, etc.), candidate reference images that match the image to be queried can be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if the texture in the image to be queried is weak or there are fewer textures, candidate reference images with higher accuracy can still be queried based on the text fields. The initial posture of the terminal obtained by performing posture solution processing based on the candidate reference images with higher accuracy is more accurate. Retrieval and precise positioning based on text field retrieval and feature matching of text area images can utilize text semantic information in the scene, thereby improving the positioning success rate in some areas with similar textures or repeated textures, and because the 2D-3D correspondence of the reference image of the text area image is utilized, the positioning accuracy is made higher.
[0485] Through the embodiments of the present disclosure, text fields can be imperceptibly integrated into visual features, so the recall rate and precision of image retrieval will be higher, and the process is imperceptible to the user, the positioning process is also more intelligent, and the user experience is better.
[0486] Through the embodiments of the present disclosure, even if there are many similar environments at different locations in the target location, and there are many similar images in different images taken at different locations in these locations, it is still possible to query candidate reference images that match the image to be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if there is interference from many similar images, the accuracy of the candidate reference images queried based on the text fields is higher, and the initial pose of the terminal obtained by performing pose solution processing based on the candidate reference images with higher accuracy is more accurate.
[0487] An exemplary embodiment of the present disclosure provides a method for determining a posture, which can be applied to a terminal and implemented by a server. Figure 17 As shown, the processing flow of the method may include the following steps:
[0488] Step S1701 : guiding the user to capture an image to be queried at a first position.
[0489] The scene at the first position includes the scene in the image to be queried.
[0490] In implementation, the first position may include any geographical location or spatial location. The scene may refer to the scene or environment in which the terminal device is located when in use, such as a room, or a field. The scene may also be the entire scene or part of the scene that can be captured by the image capture component of the terminal within a preset position range. The scene may also include the environmental background, physical objects in the environment, etc. The specific range and size of the scene are freely defined according to actual needs, and the embodiments of the present disclosure do not limit this. The scene of the first position may refer to the specific scene around the first position, and may include a certain preset geographical range or field of view. The image to be queried may be an image captured by the terminal at the first position, and the scene in the image is consistent with the physical scene. The scene in the image to be queried may be part or all of the scene corresponding to the first position. The first position is not limited to a certain precise position, and the actual position is allowed to have a certain accuracy error.
[0491] If the image to be queried is taken by the user using the terminal, some measures can be taken to ensure the image quality of the image to be queried, and the user can be guided to take the image through the user interface. In the embodiment of the present disclosure, three methods are provided to guide the user to take the image to be queried.
[0492] Optionally, step S1701 may include: capturing a first initial image; when there is no text in the first initial image, displaying or voice broadcasting a first prompt message; wherein the first prompt message is used to indicate that no text is detected in the first initial image, and prompting the user to move the position of the terminal or adjust the shooting angle of the terminal; until the terminal captures a second initial image with text at the first position, the second initial image is determined as the image to be queried.
[0493] In practice, when a user arrives at a target location, such as a shopping mall, and if it is the first time for the user to visit the mall and the user wants to view some information about the mall through the mobile phone, the user can take out the mobile phone, stand at the second position, open the camera and aim at a certain environment of the mall to shoot. At this time, the terminal can obtain the first initial image and detect whether there is text in the first initial image. When there is no text in the first initial image, such as Figure 9 As shown, the terminal can pop up a prompt box or directly broadcast the first prompt information by voice, for example, it can display "No text is detected in the current image, please try changing the position or adjusting the shooting angle". After receiving the prompt, the user moves the phone toward the location of the physical text. During the movement, the terminal continuously captures the initial image until the user reaches a suitable first position. The terminal detects the text in the second initial image and can determine the second initial image as the image to be queried and send it to the server. At this time, the terminal can also pop up a prompt box to inform the user that the text has been detected and that information query processing is being performed based on the captured image.
[0494] Optionally, step S1701 may include: capturing a third initial image; determining a text area image contained in the third initial image by performing text detection processing on the third initial image; when the text area image contained in the third initial image does not meet the preferred image condition, displaying or voice broadcasting a second prompt message; wherein the second prompt message is used to indicate that the text area image contained in the third initial image does not meet the preferred image condition, and prompting the user to move the terminal in the direction where the physical text is located; until a fourth initial image containing a text area image that meets the preferred image condition is captured at the first position, the fourth initial image is determined as the image to be queried.
[0495] Among them, the preferred image conditions may include one or more of the following conditions: the size of the text area image is greater than or equal to the size threshold, the clarity of the text area image is greater than or equal to the clarity threshold, and the texture complexity of the text area image is less than or equal to the complexity threshold.
[0496] In implementation, the terminal can detect the text area image in the third initial image and determine the size of the text area image. If the size of the text area image is small, it means that the text area image may not be very clear, and thus the current image does not meet the requirements. The terminal can also directly determine the clarity of the text area image. If the clarity is less than the clarity threshold, it means that the text area image may not be very clear, and thus the current image does not meet the requirements. The terminal can also determine the texture complexity of the text area image. If the texture complexity of the text area image is high, it means that there are more texture features in the text area image, which may interfere with the subsequent recognition of the text field in the text area image, and thus the current image does not meet the requirements. Other preferred image conditions can be reasonably set based on the preferred image conditions provided in the embodiment of the present disclosure according to actual needs. When the text area image in the third initial image does not meet one or more of the preferred image conditions, the initial image is retaken.
[0497] In actual application, when a user arrives at a target location such as a shopping mall, if it is the first time for the user to visit the mall and he wants to view some information about the mall through his mobile phone, he can take out his mobile phone, stand at the second position, open the camera and aim at a certain environment of the mall to shoot. At this time, the terminal can obtain the third initial image and detect whether the text area image contained in the third initial image meets the preferred image condition. When the size of the text area image contained in the third initial image is smaller than the size threshold, such as Figure 10 As shown, the terminal can pop up a prompt box or directly broadcast the second prompt information by voice, for example, it can display "The text box detected in the current image is small, please try to move closer to the location of the physical text". After the user receives the prompt, he moves the mobile phone toward the location of the physical text. During the movement, the terminal continuously captures the initial image until the user reaches a suitable first position. The size of the text area image detected by the terminal in the fourth initial image is greater than the size threshold. The fourth initial image can be determined as the image to be queried and sent to the server. At this time, the terminal can also pop up a prompt box to prompt the user that the current image has met the requirements and is performing information query processing based on the captured image. When the clarity and texture complexity of the text area image contained in the third initial image do not meet the requirements, the user can also be prompted based on the above method to guide the user to capture the image to be queried with higher image quality.
[0498] Optionally, step S1701 may include: capturing a fifth initial image; determining N text fields contained in the fifth initial image; obtaining M text fields contained in a reference query image, wherein the time interval between the acquisition time of the reference query image and the acquisition time of the fifth initial image is less than a duration threshold, and M is greater than or equal to 1; when any text field contained in the fifth initial image is inconsistent with each of the M text fields, displaying or voice broadcasting a third prompt message; wherein the third prompt message is used to indicate that an erroneous text field is identified in the fifth initial image, and prompting the user to move the terminal or adjust the shooting angle of the terminal; until each text field contained in the sixth initial image captured at the first position belongs to the M text fields, the sixth initial image is determined as the image to be queried.
[0499] In practice, if the fifth initial image contains one text field, a search can be performed to determine whether the text field exists among the M text fields contained in the reference query image. If the text field exists among the M text fields contained in the reference query image, the text field identified in the fifth initial image is the correct text field. If the fifth initial image contains at least two text fields, each text field contained in the fifth initial image can be acquired sequentially. Whenever a text field is acquired, a search is performed to determine whether the currently acquired text field exists among the M text fields contained in the reference query image. If the currently acquired text field exists among the M text fields contained in the reference query image, the currently acquired text field is the correct text field. Since the fifth initial image contains at least two text fields, the above determination needs to be performed multiple times. As long as any text field contained in the fifth initial image does not exist among the M text fields contained in the reference query image, the initial image can be retaken.
[0500] In actual applications, when a user comes to a target place such as a shopping mall, if the user is visiting the mall for the first time and wants to view some information about the mall through a mobile phone, he can take out the mobile phone, stand at the second position, turn on the camera and aim at a certain environment of the mall to shoot. At this time, the terminal can obtain the fifth initial image, and the terminal can detect N text area images in the fifth initial image, and respectively identify the text fields contained in each text area image. Then, the terminal can also judge the accuracy of the text field identified from the fifth initial image based on the upper and lower video frames of the fifth initial image in the video stream. The time interval between the upper and lower video frames and the fifth initial image is less than the duration threshold. For example, taking the example of only one text field being detected in the image, the text fields identified in the previous video and the next video frame of the fifth initial image are both "A35", and the text field identified in the fifth initial image is "A36", which means that the text field identified in the fifth initial image may be an incorrect text field, and then as Figure 11 As shown, a prompt box may pop up or the third prompt information may be directly broadcast by voice, for example, it may be displayed that "the wrong text may have been recognized, please try changing the position or adjusting the shooting angle". After the user receives the prompt, the mobile phone moves toward the location of the physical text. During the movement, the terminal continuously captures the initial image until the user reaches a suitable first position. When each text field recognized by the terminal in the sixth initial image belongs to the text field that can be recognized in the upper and lower video frames, the sixth initial image can be determined as the image to be queried and sent to the server. At this time, the terminal can also pop up a prompt box to prompt the user that the correct text has been detected and that information query processing is being performed based on the captured image.
[0501] Through the above methods, after each user performs an operation, the terminal can evaluate the user's operation based on preset logic and provide appropriate guidance to guide the user to capture the query image with high image quality. The three methods provided above can be used individually, in combination with any two or all three. In addition, other methods for guiding the user to capture high-quality query images can also be used together.
[0502] In addition to the above methods, if the query image is captured automatically by the terminal, the terminal's image capture component can capture the target location's video stream in real time while the terminal is moving, and extract the query image from the video stream. The query image can be a specific frame in the video stream. The terminal can also capture the target location's environment image every time a preset capture cycle is reached, using the complete environment image as the query image.
[0503] The terminal captures different query images at different positions, and the terminal's initial position and posture can be determined based on the features of the captured query image, where the initial position and posture include the terminal's current initial position and posture.
[0504] Step S1702: Determine N text region images contained in the query image.
[0505] Wherein, N is greater than or equal to 1.
[0506] In implementation, after the query image is acquired, text detection processing can be performed to determine N text region images in the query image.
[0507] Step S1703: Identify the text fields included in each text region image.
[0508] In implementation, the text contained in each text area image can be identified separately to obtain N text fields. A text field can be one character or multiple characters. For example, a text field can be "A", "3", or "5", or a text field can be "A35". "A", "3", and "5" each represent one character, and "A35" represents three characters. One character or multiple characters can serve as a text field. In one possible implementation, the characters contained in the continuous area image portion in the image to be queried can serve as a text field. The text field can be stored in the terminal in the form of a character string.
[0509] Optionally, the N text fields detected in the query image can be screened to further extract text fields of a significant type. A significant type of text field is an identifying field that can clearly or uniquely identify an environment. The terminal can input the N text fields contained in the query image into a pre-trained text classifier to obtain the text type of each text field contained in the query image, and determine that the text type is a preset significant type of text field. This execution logic can also be set to be completed on the server, that is, the terminal sends all text fields to the server, and the server screens out the significant type of text fields from the N text fields based on similar logic. If the terminal screens out the significant type of text field from the N text fields, then what the terminal ultimately sends to the server includes the query image and the significant type of text field.
[0510] In implementation, text fields and query images based on salient types are more conducive to improving the accuracy of the determined initial pose.
[0511] Step S1704: Obtain the location area of the text area image in the image to be queried, and send N text fields, the image to be queried, and the location area to the server.
[0512] In implementation, the location region may be the location coordinates of the text box corresponding to the text area image. If the text box is a square, the location region may be the location coordinates of the four corners of the text box, or the location coordinates of two diagonal corners. If the text box is a circle, the location region may be the location coordinates of the circle's center and the circle's radius.
[0513] The server can determine the terminal's initial pose based on the N text fields, the query image, and the location region sent by the terminal, and then return the initial pose to the terminal. The specific method by which the server determines the terminal's initial pose based on the N text fields, the query image, and the location region sent by the terminal will be described later.
[0514] Step S1705: receiving the initial posture of the terminal at the first position returned by the server.
[0515] In implementation, the terminal can perform navigation, route planning, obstacle avoidance, etc. based on the received initial position at the first location. The initial position at the first location is determined by the server in real time based on the N text fields, the image to be queried, and the location area sent by the terminal.
[0516] Optionally, after receiving the initial posture returned by the server, the method provided by the embodiment of the present disclosure may further include: determining the real-time posture according to the initial posture and the posture change of the terminal.
[0517] In implementation, if the initial posture is determined by a query image in the video, the posture change of the terminal can be determined through real-time positioning and SLAM tracking technology. The real-time posture is determined based on the initial posture and the posture change of the terminal.
[0518] Optionally, in addition to performing navigation, route planning, obstacle avoidance and other processing based on the real-time posture, after receiving the initial posture returned by the server, the method provided by the embodiment of the present disclosure may also include: obtaining a preview stream of the current scene; determining the preset media content contained in the digital map corresponding to the scene in the preview stream based on the real-time posture; and rendering the media content in the preview stream.
[0519] In implementation, if the terminal is a mobile phone or wearable AR device, a virtual scene can be constructed based on the real-time pose. First, the terminal can obtain a preview stream of the current scene. For example, a user can capture a preview stream of the current environment in a shopping mall. Next, the terminal can determine the real-time pose using the method described above. Subsequently, the terminal can obtain a digital map, which records the three-dimensional coordinates of various locations in the world coordinate system. Preset media content exists at these preset three-dimensional coordinates. The terminal can determine the target three-dimensional coordinates corresponding to the real-time pose in the digital map. If the target three-dimensional coordinates contain corresponding preset media content, the preset media content is obtained. For example, if a user is photographing a target store, the terminal recognizes the real-time pose and determines that the camera is currently facing the target store. The preset media content corresponding to the target store can be obtained. The preset media content corresponding to the target store can include information about the target store, such as which products are worth purchasing. Based on this, the terminal can render the media content in the preview stream. At this point, the user can view the preset media content corresponding to the target store in a preset area near the image of the target store on the phone. After the user has viewed the preset media content corresponding to the target store, he or she can have a general understanding of the target store.
[0520] Different digital maps can be set for different places, so that when the user moves to other places, he can also obtain the preset media content corresponding to the real-time posture based on the method of rendering media content provided in the embodiment of the present disclosure, and render the media content in the preview stream.
[0521] The following describes a specific method in which the server determines the initial position of the terminal based on the N text fields, the image to be queried, and the location area sent by the terminal.
[0522] The server can receive N text fields, the image to be queried and the location area sent by the terminal, and determine the candidate reference image according to the N text fields based on the correspondence between the pre-stored reference image and the text field. The server can also determine the target text area image in the image to be queried based on the location area, and the image to be queried can be the complete environment image taken by the terminal at this time. The text area image included in the candidate reference image is obtained, and the candidate reference image can be the candidate environment image in the complete environment image taken in advance. Then, the initial posture of the terminal can be determined based on the target text area image and the text area image included in the candidate reference image, and the determined initial posture is sent to the terminal. Wherein, N is greater than or equal to 1. The image to be queried is obtained based on the image captured by the terminal at the first position, and the scene at the first position includes the scene in the image to be queried.
[0523] During implementation, through an offline calibration process, the server can pre-establish a database that can store pre-taken environmental images at various locations in the target location. Through the offline calibration process, the server can also pre-determine text area images in each pre-taken environmental image and obtain a pre-established 2D-3D correspondence. The 2D-3D correspondence includes a large number of 3D points and corresponding 2D points in the environmental image. Each 3D point corresponds to a physical point near the location of the physical text in the target location, and each 3D point corresponds to the three-dimensional position information of a corresponding physical point in real space. In addition, the server can pre-identify text fields in the text area images in each pre-taken environmental image and store the corresponding text area images and text fields.
[0524] In order to improve the search speed of the text field, a search index may be established based on the text field in the database, and the pre-identified text area images, text fields and search indexes may be stored in correspondence.
[0525] The server can perform retrieval processing on the N text fields sent by the terminal based on the search index to determine the candidate reference images corresponding to the N text fields. In the process of determining the candidate reference images corresponding to the N text fields, each text field can be acquired one by one. Whenever a text field is acquired, the candidate reference image corresponding to the currently acquired text field is determined. In this way, the candidate reference image corresponding to each text field can be determined one by one. Next, the text area image included in the candidate reference image can be acquired, and based on the text area image included in the candidate reference image and the target text area image, the initial posture of the terminal at the first position is determined.
[0526] Optionally, based on the target text area image and the text area image in the candidate reference image, the step of determining the initial posture of the terminal at the first position may include: determining the target reference image in the candidate reference image; wherein the scene at the first position includes the scene in the target reference image; and determining the initial posture of the terminal at the first position based on the target text area image and the text area image in the target reference image.
[0527] During implementation, the candidate reference images can be screened to determine the target reference image. The method provided in the embodiments of this disclosure introduces three methods for screening candidate reference images, which will be described in detail later. When the number of candidate reference images is equal to one, the candidate reference image can be directly determined as the target reference image. After determining the target reference image, the initial pose of the terminal can be determined based on the query image and the target reference image.
[0528] Optionally, the step of determining the initial posture of the terminal at the first position based on the image to be queried and the target reference image may include: determining the 2D-2D correspondence between the image to be queried and the target reference image; and determining the initial posture of the terminal at the first position based on the 2D-2D correspondence and the 2D-3D correspondence of the preset target reference image.
[0529] In an implementation, the 2D-2D correspondence between the query image and the target reference image may include a 2D-2D correspondence between the target text region image and the text region image contained in the target reference image. The 2D-3D correspondence of the target reference image may include a 2D-3D correspondence between the text region image contained in the target reference image.
[0530] After determining the 2D-2D correspondence between the target text area image and the text area image contained in the target reference image, the 2D-2D correspondence and the 2D-3D correspondence of the text area image contained in the target reference image can be input into the posture estimation module. The posture estimation module can solve the initial posture of the terminal and output the initial posture of the terminal at the first position.
[0531] Optionally, before determining the 2D-2D correspondence between the target text area image and the text area image contained in the target reference image, the target text area image may be subjected to image enhancement processing to obtain the target text area image after image enhancement processing, and then, the 2D-2D correspondence between the target text area image after image enhancement processing and the text area image contained in the target reference image may be determined.
[0532] The following describes four methods for screening candidate reference images in the method provided by the embodiments of the present disclosure.
[0533] Optionally, the step of determining the target reference image from the candidate reference images may include: determining the image similarity between each candidate reference image and the query image; and determining the candidate reference image whose image similarity is greater than or equal to a preset similarity threshold as the target reference image.
[0534] Optionally, the step of determining a target reference image among candidate reference images may include: obtaining global image features of each candidate reference image; determining global image features of the image to be queried; determining the distance between the global image features of each candidate reference image and the global image features of the image to be queried; and determining a candidate reference image whose distance is less than or equal to a preset distance threshold as a target reference image.
[0535] Optionally, the candidate reference images may be screened based on the terminal's positioning information. The method provided by the disclosed embodiment may further include: the terminal obtaining the terminal's positioning information; and sending the positioning information to the server. Accordingly, the step of determining the target reference image from the candidate reference images in the server may include: receiving the positioning information sent by the terminal; obtaining the shooting positions corresponding to each candidate reference image; and determining, from each candidate reference image, the target reference image whose shooting position matches the positioning information.
[0536] Optionally, when the number of text fields contained in the query image is greater than 1, the step of determining the target reference image from the candidate reference images may include: determining a target reference image containing N text fields from each candidate reference image.
[0537] Optionally, based on the correspondence between pre-stored reference images and text fields, the step of determining candidate reference images according to N text fields may include: inputting the N text fields contained in the image to be queried into a pre-trained text classifier to obtain the text type of each text field contained in the image to be queried; determining a text field whose text type is a preset significant type; and based on the correspondence between pre-stored reference images and text fields, searching for candidate reference images corresponding to the text field of the significant type.
[0538] Optionally, based on the target text area image and the text area image contained in the target reference image, the step of determining the initial posture of the terminal may include: for each text area image contained in the target reference image, determining the initial correspondence between the image key points in the target text area image and the image key points in the text area image contained in the current target reference image, performing geometric verification processing on each pair of image key points in the initial correspondence, excluding image key points with incorrect matches in the initial correspondence, and obtaining the target correspondence; if the target correspondence contains pairs of image key points greater than or equal to a preset threshold, it indicates that the target reference image and the image to be queried are images captured near the first position; based on the image to be queried and the target reference image, the initial posture of the terminal is determined.
[0539] Based on the above, if Figure 18 FIG. 1 is a flow chart of a method for determining a posture provided by an embodiment of the present disclosure. The flow of the method for determining a posture may include:
[0540] Step S1801: Capture video stream.
[0541] Step S1802: extract the image to be queried from the video stream.
[0542] Step S1803: Perform text detection on the query image.
[0543] Step S1804: determine whether a text box is detected.
[0544] Step S1805: If no text box is detected, the user is guided to adopt a better shooting method to shoot.
[0545] Step S1806: If a text box is detected, the characters in the text box are recognized.
[0546] Step S1807: judging whether the character corresponding to the query image is correctly recognized by the characters corresponding to the upper and lower video frames of the query image. If not, proceeding to step S1805.
[0547] Step S1808: If the character recognition result corresponding to the query image is correct, image enhancement processing is performed on the text area image.
[0548] Step S1809: extracting image key points of the text region image after image enhancement processing.
[0549] Step S1810: If the character recognition result corresponding to the image to be queried is correct, an image search is performed based on the character recognition result to obtain the target environment image.
[0550] Step S1811 , performing key point matching based on the image key points of the text region image after image enhancement processing and the image key points of the text region image of the target environment image.
[0551] Step S1812: establishing a target 2D-3D correspondence based on the key point matching results.
[0552] Step S1813: Perform pose estimation based on the target 2D-3D correspondence.
[0553] The processing of some terminals and the processing of the server in the embodiments disclosed herein are similar to the processing of some terminals and the processing of the server in the previously disclosed embodiments. For some of the shared parts, no further description is made in the embodiments disclosed herein. You may refer to the description of the processing of the terminals and the processing of the server in the previously disclosed embodiments. It should be noted that the processing of each step in S1801-S1813 can be performed by the terminal or by the server, and the possible interaction combinations are various and are not listed here one by one. In any of the above possible interaction combinations, when implementing the above steps based on the concept of the present invention, those skilled in the art can construct the communication process between the server and the terminal, such as what necessary information is interacted and transmitted between the server and the terminal; these are not exhaustively listed or detailed in the present invention.
[0554] Through the embodiments of the present disclosure, in some scenes with weak textures or high texture similarity (such as in corridors or where the wall occupies a large area of the image to be queried, etc.), candidate reference images that match the image to be queried can be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if the texture in the image to be queried is weak or there are fewer textures, candidate reference images with higher accuracy can still be queried based on the text fields. The initial posture of the terminal obtained by performing posture solution processing based on the candidate reference images with higher accuracy is more accurate. Retrieval and precise positioning based on text field retrieval and feature matching of text area images can utilize text semantic information in the scene, thereby improving the positioning success rate in some areas with similar textures or repeated textures, and because the 2D-3D correspondence of the reference image of the text area image is utilized, the positioning accuracy is made higher.
[0555] Through the embodiments of the present disclosure, text fields can be imperceptibly integrated into visual features, so the recall rate and precision of image retrieval will be higher, and the process is imperceptible to the user, the positioning process is also more intelligent, and the user experience is better.
[0556] Through the embodiments of the present disclosure, even if there are many similar environments at different locations in the target location, and there are many similar images in different images taken at different locations in these locations, it is still possible to query candidate reference images that match the image to be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if there is interference from many similar images, the accuracy of the candidate reference images queried based on the text fields is higher, and the initial pose of the terminal obtained by performing pose solution processing based on the candidate reference images with higher accuracy is more accurate.
[0557] An exemplary embodiment of the present disclosure provides a method for determining a posture, which can be applied to a terminal, such as Figure 19 As shown, the processing flow of the method may include the following steps:
[0558] Step S1901: The terminal obtains an image to be queried at a first location.
[0559] The scene at the first position includes the scene in the image to be queried.
[0560] Step S1902: Send the image to be queried to the server, so that the server determines N text fields contained in the image to be queried, and determines the initial posture of the terminal at the first position according to the N text fields and the image to be queried.
[0561] Wherein, N is greater than or equal to 1;
[0562] Step S1903: Receive the initial posture of the terminal at the first position returned by the server.
[0563] In one possible implementation, the method also includes: obtaining the positioning information of the terminal; sending the positioning information to the server; determining the initial posture of the terminal at the first position based on N text fields and the image to be queried, including: determining the initial posture of the terminal at the first position based on N text fields, the image to be queried and the positioning information.
[0564] In a possible implementation, after receiving the initial posture returned by the server, the method provided by the embodiment of the present disclosure may further include: determining a real-time posture according to the initial posture and the posture change of the terminal.
[0565] In one possible implementation, after receiving the initial posture returned by the server, the method provided by the embodiment of the present disclosure may further include: obtaining a preview stream of the current scene; determining, based on the real-time posture, preset media content contained in the digital map corresponding to the scene in the preview stream; and rendering the media content in the preview stream.
[0566] Through the embodiments of the present disclosure, even if there are many similar environments at different locations in the target location, and there are many similar images in different images taken at different locations in these locations, it is still possible to query candidate reference images that match the image to be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if there is interference from many similar images, the accuracy of the candidate reference images queried based on the text fields is higher, and the initial pose of the terminal obtained by performing pose solution processing based on the candidate reference images with higher accuracy is more accurate.
[0567] An exemplary embodiment of the present disclosure provides a method for determining a posture, which can be applied to a server and implemented in conjunction with a terminal, such as Figure 20 As shown, the processing flow of the method may include the following steps:
[0568] Step S2001: receiving an image to be queried sent by a terminal.
[0569] The image to be queried is obtained based on an image captured by the terminal at a first position, and the scene at the first position includes the scene in the image to be queried.
[0570] The first position may include any geographical location or spatial location. The scene may refer to the scene or environment in which the terminal device is located when in use, such as a room, or a field. The scene may also be the entire scene or part of the scene that can be captured by the image capture component of the terminal within a preset position range. The scene may also include the environmental background, physical objects in the environment, etc. The specific range and size of the scene are freely defined according to actual needs, and the embodiments of the present disclosure do not limit this. The scene of the first position may refer to the specific scene around the first position, and may include a certain preset geographical range or field of view. The image to be queried may be an image captured by the terminal at the first position, and the scene in the image is consistent with the physical scene. The scene in the image to be queried may be part or all of the scene corresponding to the first position. The first position is not limited to a certain precise position, and the actual position is allowed to have a certain accuracy error.
[0571] In implementation, a user may use a terminal to capture an image to be queried, and send the image to be queried to a server, and the server may receive the image to be queried sent by the terminal.
[0572] In order to ensure the quality of the query image received by the server, some means can be used to ensure the image quality of the captured query image. For example, a prompt message can be displayed in the user interface, which is used to prompt the user to capture the query image containing text.
[0573] Optionally, if the terminal has a text detection function, the three methods of guiding the user to shoot the image to be queried provided in the above step S801 may also be used to ensure the quality of the image to be queried.
[0574] In addition, when the terminal captures the image to be queried, it can also obtain the terminal's positioning information and send the terminal's positioning information to the server, and the server can receive the positioning information sent by the terminal.
[0575] Step S2002: Determine N text fields contained in the image to be queried.
[0576] Wherein, N is greater than or equal to 1.
[0577] In implementation, after obtaining the image to be queried, N text fields contained in the image to be queried can be identified based on OCR technology. Specifically, text detection processing can be performed to determine N text area images in the image to be queried, and the text contained in each text area image can be identified to obtain N text fields. A text field can be 1 character or multiple characters. For example, a text field can be "A", "3", or "5", and a text field can also be "A35". "A", "3", and "5" each represent 1 character, and "A35" represents 3 characters. 1 character or multiple characters can be used as a text field. In a possible implementation method, the characters contained in a continuous area image portion in the image to be queried can be used as a text field. The text field can be stored in the server in the form of a character string.
[0578] In the process of identifying the text fields contained in the query image, the server can first perform text detection on the query image and output the text box position coordinates corresponding to the text area image contained in the query image. In the embodiment of the present disclosure, the text detection process can be performed based on the deep learning target detection algorithm (Single Shot Multi Box Detector, SSD). Multiple text boxes can be detected in a query image, and each text box corresponds to one text field.
[0579] Assuming the target location is an underground garage, parking area signs can be set on the pillars of the underground garage. Figure 12 As shown in the figure, the current parking area mark "A35" is set on the pillar of the underground garage. When the terminal collects the query image in the current parking area, the collected query image is likely to include the current parking area mark "A35". When the server performs text detection on the query image including the current parking area mark "A35", it can output the position coordinates of the text box corresponding to the "A35" area image. Or, as Figure 13 As shown, when the terminal collects the image to be queried in the corridor of a building, the collected image to be queried is likely to include the current floor identification "3B" or "3E". When the server performs text detection processing on the image to be queried including the current floor identification "3B" or "3E", it can output the text box position coordinates corresponding to the "3B" or "3E" area image.
[0580] In the embodiment of the present disclosure, Figure 14As shown, after determining the text box position coordinates corresponding to the text region image contained in the query image, the text region image can be extracted from the query image based on the text box position coordinates. Convolutional Neural Network (CNN) feature extraction is performed on the text region image, and the extracted CNN features are then input into a recurrent neural network (Long Short-Term Memory, LSTM) for encoding. The encoded CNN features are then classified, ultimately outputting the text field in the text region image, such as "A35."
[0581] Optionally, the N text fields detected in the query image may be screened to further extract text fields of a salient type. A salient type of text field is an identifying field that can clearly or uniquely identify an environment. The server may input the N text fields contained in the query image into a pre-trained text classifier to obtain the text type of each text field contained in the query image, and determine that the text type is a text field of a preset salient type.
[0582] In practice, the target location may contain a large number of text fields, some of which are helpful in identifying the current environment, while some text fields may interfere with the process of identifying the current environment. Text fields that are helpful in identifying the current environment can be used as significant types of text fields. Effective text field collection rules can be formulated in advance. In application, text fields with identification in the target location can be selected as positive samples, such as parking area signs "A32" and "B405" in underground garages. At the same time, text fields that are not identification in the target location can also be selected as negative samples, and the classifier can be trained based on positive and negative samples.
[0583] After extracting N text fields from the query image, each of these N text fields can be fed into a trained classifier. If the classifier outputs a value close to or equal to 1, the current text field is considered salient; if the output value is close to or equal to 0, the current text field is considered non-salient. This helps improve the accuracy of the initial pose determined based on the salient text fields and the query image.
[0584] Step S2003 : Based on the pre-stored correspondence between reference images and text fields, candidate reference images are determined according to the N text fields.
[0585] During implementation, through the offline calibration process, the server can pre-establish a database in which reference images taken at various locations in the target location can be stored. Through the offline calibration process, the server can also obtain a pre-established 2D-3D correspondence relationship. The 2D-3D correspondence relationship contains a large number of 3D points and corresponding 2D points in the reference image. Each 3D point corresponds to a physical point in the target location, and each 3D point corresponds to the three-dimensional position information of a corresponding physical point in the actual space. In addition, the server can pre-identify the text field in each reference image and store the corresponding reference image and text field. When the image to be queried is a complete environmental image taken by the terminal, the reference image can be a complete environmental image pre-taken at various locations in the target location.
[0586] In order to improve the search speed of the text field, a search index (Global Index) may be established based on the text field in the database, and each reference image, text field, and search index may be stored in correspondence.
[0587] After the server determines the N text fields in the query image, it can perform a search process on the N text fields contained in the query image based on the search index to determine candidate reference images corresponding to the N text fields. The candidate reference image can be a target environment image from a pre-captured environment image. In determining the candidate reference images corresponding to the N text fields, each text field can be acquired one by one. Each time a text field is acquired, a candidate reference image corresponding to the currently acquired text field is determined. In this way, the candidate reference image corresponding to each text field can be determined one by one.
[0588] For an image in which two or more text fields in a candidate reference image are identical to the text fields contained in the image to be queried, this type of candidate reference image will be determined twice or multiple times based on the above-mentioned determination method, that is, the first candidate reference image can be determined based on the first text field in the image to be queried, and the first candidate reference image can also be determined based on the second text field in the image to be queried. Therefore, the first candidate reference image includes both the first text field and the second text field. Therefore, a deduplication operation can be performed on all determined candidate reference images to remove candidate reference images that have been determined multiple times. After determining the candidate reference images, the server can determine the initial posture of the terminal based on the image to be queried and the candidate reference images, and send the determined initial posture to the terminal.
[0589] Optionally, the above description is about the case where the terminal sends only one image to be queried to the server. If the terminal sends multiple images to be queried to the server, an initial position of the terminal can be determined based on each image to be queried, and multiple initial positions can be obtained. Among the multiple initial positions, the target initial position is determined based on probability statistics, and the target initial position is sent to the terminal. Among the multiple initial positions, the step of determining the target initial position based on probability statistics may include: determining the target initial position that appears the most times among the multiple initial positions.
[0590] Optionally, based on the image to be queried and the candidate reference images, the step of determining the initial posture of the terminal may include: determining a target reference image in the candidate reference images; wherein the scene at the first position includes the scene in the target reference image; and determining the initial posture of the terminal based on the image to be queried and the target reference image.
[0591] In implementation, since some of the candidate reference images are interference images, that is, they are not necessarily images pre-taken near the first position where the image to be queried is taken, but the text field corresponding to the candidate reference image and the text field corresponding to the image to be queried happen to be consistent, the interference image is also used as a candidate reference image to determine the initial posture of the terminal, which will affect the accuracy of the initial posture. Therefore, the candidate reference images can be screened to determine the target reference image. The method provided in the embodiment of the present disclosure introduces four ways of screening candidate reference images, which will be introduced in detail later. When the number of candidate reference images is equal to 1, the candidate reference image can be directly determined as the target reference image. After the target reference image is determined, the initial posture of the terminal can be determined based on the image to be queried and the target reference image.
[0592] Optionally, the step of determining the initial posture of the terminal based on the image to be queried and the target reference image may include: determining the 2D-2D correspondence between the image to be queried and the target reference image; and determining the initial posture of the terminal based on the 2D-2D correspondence and the 2D-3D correspondence of the preset target reference image.
[0593] In implementation, during the offline calibration process, the server can pre-extract local image features of each reference image, and store the corresponding environmental images, local image features, and the aforementioned text fields. Local image features can include image key points, such as corner points and other characteristic pixel points in other images. The server can determine the local image features corresponding to the target reference image based on the correspondence between the reference image and the local image features. After the server receives the image to be queried sent by the terminal, it can extract the local image features of the image to be queried, and perform feature matching on the local image features of the image to be queried and the local image features corresponding to the target reference image, that is, match the image key points corresponding to the image to be queried and the image key points corresponding to the target reference image. The image key points are 2D points, and a 2D-2D correspondence between the image to be queried and the target reference image can be obtained. The 2D-2D correspondence between the image to be queried and the target reference image can include a 2D-2D correspondence between the complete environmental image captured by the terminal and the target environmental image in the pre-captured complete environmental image.
[0594] For example, if there are three image key points in the query image (including A1, B1, and C1), and five image key points in the target reference image (including A2, B2, C2, D2, and E2), feature matching can determine three sets of correspondences: A1-B2, B1-E2, and C1-A2. Of course, the feature matching process in actual applications is much more complex, involving a greater number of image key points. Here, we only use a few image key points as examples. It should be noted that, in theory, the image key points corresponding to the query image and the image key points corresponding to the target environment image should correspond to the same physical points.
[0595] During the offline calibration process, the server can establish a 3D point cloud of the target location based on the reference images taken at various locations in the target location. Each pixel in each reference image corresponds to a 3D point in the 3D point cloud, which can be recorded as an initial 2D-3D correspondence. After the server determines the image key points of each reference image, it can determine the 3D point in the 3D point cloud corresponding to each image key point in the target reference image based on the initial 2D-3D correspondence, which can be recorded as the 2D-3D correspondence of the target reference image. The 2D-3D correspondence of the target reference image can be the 2D-3D correspondence of the target reference image corresponding to the first position of the terminal when capturing the image to be queried. During the online positioning process, the server can determine the 2D-2D correspondence between the query image and the target reference image, that is, the correspondence between the image key points of the query image and the image key points of the target environment image. Then, the server can determine the 3D points corresponding to the image key points of the query image in the 3D point cloud based on the 2D-2D correspondence between the query image and the target reference image, as well as the 2D-3D correspondence of the target reference image.
[0596] After determining the 3D points corresponding to the image key points of the query image in the 3D point cloud, the 3D points corresponding to the image key points of the query image in the 3D point cloud, the position information of each image key point of the query image and the three-dimensional position information of each corresponding 3D point can be input into the posture estimation module. The posture estimation module can solve the posture of the terminal and output the initial posture of the terminal.
[0597] The following describes four methods for screening candidate reference images in the method provided by the embodiments of the present disclosure.
[0598] Optionally, the step of determining the target reference image from the candidate reference images may include: determining the image similarity between each candidate reference image and the query image; and determining the candidate reference image whose image similarity is greater than or equal to a preset similarity threshold as the target reference image.
[0599] During implementation, the server may calculate the image similarity between each candidate reference image and the query image based on a preset image similarity algorithm. Subsequently, among the candidate reference images, images whose image similarity is greater than a preset similarity threshold may be determined as target reference images. The number of target reference images determined may be one or more. Alternatively, the candidate reference images may be sorted in descending order of image similarity, and a preset number of images ranked at the top of the list may be determined as target reference images. The image similarity algorithm may include a K-nearest neighbor algorithm, for example. The preset similarity threshold may be determined based on experience and may be set to a more reasonable value based on experience.
[0600] Optionally, the step of determining a target reference image among candidate reference images may include: obtaining global image features of each candidate reference image; determining global image features of the image to be queried; determining the distance between the global image features of each candidate reference image and the global image features of the image to be queried; and determining a candidate reference image whose distance is less than or equal to a preset distance threshold as a target reference image.
[0601] In implementation, the server can determine the global image features corresponding to each candidate reference image based on the correspondence between the pre-stored reference images and the global image features, and can also extract the global image features corresponding to the image to be queried. The global image features can be data represented in the form of vectors, so the distance between the global image features corresponding to each candidate reference image and the global image features corresponding to the image to be queried can be calculated. The distance can be a distance of the Euclidean distance or other types. After calculating the distance, the candidate reference image whose distance is less than or equal to the preset distance threshold can be determined as the target reference image. The preset distance threshold can be determined based on experience, and can be set to a more reasonable value based on experience. Alternatively, after calculating the distance, the candidate reference images can be sorted in order of distance from small to large, and a preset number of candidate reference images ranked first are selected as the target reference images.
[0602] like Figure 14 As shown, in the method provided in the embodiment of the present disclosure, a VGG network can be used to extract global image features. The environment image can be input into the VGG network, which can perform CNN feature extraction on the environment image. The VGG network includes multiple network layers, and the output of the second-to-last fully connected layer in the multiple network layers can be selected as the extracted CNN features. The extracted CNN features are then normalized using the L2 normalization method to obtain 4096-dimensional normalized features, which are the global image features of the environment image. In practical applications, global image features of the environment image can also be extracted using other methods, which are not limited by the embodiment of the present disclosure.
[0603] like Figure 15As shown, the system provided by the embodiments of the present disclosure may include: a video stream input module 1501, an image extraction module 1502, a text box detection module 1503, a text recognition module 1504, a global feature extraction module 1505, a local feature extraction module 1506, an image retrieval module 1507, a 2D-2D feature matching module 1508, a 2D-3D matching module 1509, and a pose estimation module 1510. The video stream input module 1501 can be used to acquire a video stream, the image extraction module 1502 can be used to extract a video frame from the video stream, the text box detection module 1503 can be used to detect a text region image in the video frame, and the text recognition module 1504 can be used to determine the text field in the text region image. The global feature extraction module 1505 can be used to extract global image features of the video frame, and the local feature extraction module 1506 can be used to extract local image features of the video frame, such as image key points. The operations in the global feature extraction module 1505 and the local feature extraction module 1506 can be performed in parallel. The image retrieval module 1507 can be used to search for a target reference image based on the text field and the global image features of the video frame. The 2D-2D feature matching module 1508 can be used to determine the 2D-2D correspondence between the video frame and the target reference image based on the local image features of the video frame and the local image features of the target reference image. The 2D-3D matching module 1509 can be used to determine the 2D-3D correspondence of the video frame based on the 2D-2D correspondence between the video frame and the target reference image. The pose estimation module 1510 can be used to determine the initial pose based on the 2D-3D correspondence of the video frame.
[0604] The video stream input module 1501 and image extraction module 1502 can be deployed on a terminal in the system. The text box detection module 1503, text recognition module 1504, global feature extraction module 1505, local feature extraction module 1506, image retrieval module 1507, 2D-2D feature matching module 1508, 2D-3D matching module 1509, and pose estimation module 1510 can be deployed on a server in the system. The video stream input module 1501 and image extraction module 1502 can be implemented using the acquisition module 2701 of the terminal-side device, while the text box detection module 1503, text recognition module 1504, global feature extraction module 1505, local feature extraction module 1506, image retrieval module 1507, 2D-2D feature matching module 1508, 2D-3D matching module 1509, and pose estimation module 1510 can be implemented using the determination module 2802 of the server-side device.
[0605] Optionally, the candidate reference images may be screened based on the terminal's positioning information. The method provided by the disclosed embodiment may further include: receiving positioning information sent by the terminal; obtaining the shooting positions corresponding to each candidate reference image; and determining, among the candidate reference images, a target reference image whose shooting position matches the positioning information.
[0606] During implementation, the server can divide the target location into preset unit areas, for example, it can be divided into units of 100m×100m to obtain multiple sub-areas. In the process of dividing the sub-areas, the boundaries of adjacent sub-areas can be overlapped to a certain extent. The server can calibrate the sub-area to which each reference image belongs. During the online positioning process, the terminal can first collect the current positioning information of the terminal based on the global positioning system GPS or the mobile location service LBS, and send the positioning information to the server. The server can determine the target sub-area to which the positioning belongs. Then, among the candidate reference images, the target reference image whose shooting position belongs to the same target sub-area can be determined, that is, the target reference image whose shooting position matches the positioning information.
[0607] For example, the third floors of two adjacent buildings both have the sign "301", but there is a certain distance between the two adjacent buildings. Even if the signs are repeated, through positioning, the terminal's position can be first circled within a certain range, and then the target reference image that matches the sign "301" can be searched within the certain range.
[0608] Optionally, when the number of text fields contained in the query image is greater than 1, the step of determining the target reference image from the candidate reference images may include: determining a target reference image containing N text fields from each candidate reference image.
[0609] In practice, in the process of determining candidate reference images corresponding to multiple text fields, each text field can be acquired one by one. Whenever a text field is acquired, the candidate reference image corresponding to the currently acquired text field is determined. In this way, the candidate reference image corresponding to each text field can be determined one by one. In order to further screen the candidate reference images, a target reference image that includes multiple text fields in the query image can be determined in each candidate reference image. If the same target reference image includes multiple text fields contained in the query image, then the probability that the shooting position of the target reference image is close to the shooting position of the query image is very high, and the accuracy of the initial posture determined based on this target reference image is also high.
[0610] Optionally, based on the correspondence between pre-stored reference images and text fields, the step of determining candidate reference images according to N text fields may include: inputting the N text fields contained in the image to be queried into a pre-trained text classifier to obtain the text type of each text field contained in the image to be queried; determining a text field whose text type is a preset significant type; and based on the correspondence between pre-stored reference images and text fields, searching for candidate reference images corresponding to the text field of the significant type.
[0611] In practice, the target location may contain a large number of text fields. Some text fields are helpful in identifying the current environment, while others interfere with the process. Text fields that are helpful in identifying the current environment can be considered as significant types of text fields. Effective text field collection rules can be pre-established. In the application, text fields that are identifiable in the target location can be selected as positive samples, while text fields that are not identifiable in the target location can also be selected as negative samples. The classifier is then trained based on these positive and negative samples.
[0612] The algorithm for recognizing text fields and the segmentation accuracy of templates (such as AI segmentation templates) can be determined based on user needs. This means that several characters can be segmented into a text field. A single character can be considered a text field, or all characters in a continuous image area can be considered a text field, or all characters in an area can be considered a text field.
[0613] After extracting N text fields from the query image, each of these N text fields can be fed into a trained classifier. If the classifier outputs a value close to or equal to 1, the current text field is considered salient; if the output value is close to or equal to 0, the current text field is considered non-salient. This helps improve the accuracy of the initial pose determined based on the salient text fields and the query image.
[0614] Step S2004: determining an initial posture of the terminal at the first position based on the image to be queried and the candidate reference images.
[0615] In implementation, for each target reference image, the initial correspondence between the image key points in the image to be queried and the image key points in the target reference image can be determined, and each pair of image key points in the initial correspondence is geometrically verified to eliminate image key points with incorrect matches in the initial correspondence to obtain the target correspondence. If the target correspondence contains pairs of image key points greater than or equal to a preset threshold, it means that the target reference image and the image to be queried are images collected in the same environment. Based on the image to be queried and the target reference image, the initial posture of the terminal is determined.
[0616] In implementation, the initial correspondence between the image key points in the query image and the image key points in the target reference image can be determined, and the initial correspondence includes multiple pairs of image key points. Then, each pair of image key points can be geometrically verified to exclude image key points with incorrect matches in the initial correspondence. For example, the initial correspondence includes a total of 150 pairs of image key points. 30 pairs of image key points can be excluded through geometric verification. These 30 pairs of image key points are not actually matched image key points, and the target correspondence can be obtained. Finally, it can be determined whether the target correspondence contains a number of image key points greater than or equal to a preset threshold. In the embodiment of the present disclosure, taking the preset threshold of 100 as an example, after excluding 30 pairs of image key points from the 150 pairs of image key points, 120 pairs of image key points remain, which are greater than the preset threshold of 100 pairs. Therefore, the target reference image and the query image to which the 120 pairs of remaining image key points belong are images collected in the same environment. If the target correspondence contains image key points with a value less than a preset threshold, it means that the target reference image and the query image are not environmental images captured in the same environment, and thus the target reference image can be omitted from determining the initial pose.
[0617] Step S2005: Send the initial posture to the terminal.
[0618] In implementation, the terminal receives the initial pose sent by the server.
[0619] After receiving the initial position, the terminal can perform navigation, route planning, obstacle avoidance and other processing based on the received initial position.
[0620] Optionally, after receiving the initial posture returned by the server, the method provided by the embodiment of the present disclosure may further include: the terminal determining the real-time posture according to the initial posture and the posture change of the terminal.
[0621] In implementation, if the initial posture is determined by a query image in the video, the posture change of the terminal can be determined through real-time positioning and SLAM tracking technology. The real-time posture is determined based on the initial posture and the posture change of the terminal.
[0622] Optionally, in addition to performing navigation, route planning, obstacle avoidance and other processing based on the real-time posture, after receiving the initial posture returned by the server, the method provided by the embodiment of the present disclosure may also include: the terminal obtains a preview stream of the current scene; based on the real-time posture, determines the preset media content contained in the digital map corresponding to the scene in the preview stream; and renders the media content in the preview stream.
[0623] In implementation, if the terminal is a mobile phone or wearable AR device, a virtual scene can be constructed based on the real-time pose. First, the terminal can obtain a preview stream of the current scene. For example, a user can capture a preview stream of the current environment in a shopping mall. Next, the terminal can determine the real-time pose using the method described above. Subsequently, the terminal can obtain a digital map, which records the three-dimensional coordinates of various locations in the world coordinate system. Preset media content exists at these preset three-dimensional coordinates. The terminal can determine the target three-dimensional coordinates corresponding to the real-time pose in the digital map. If the target three-dimensional coordinates contain corresponding preset media content, the preset media content is obtained. For example, if a user is photographing a target store, the terminal recognizes the real-time pose and determines that the camera is currently facing the target store. The preset media content corresponding to the target store can be obtained. The preset media content corresponding to the target store can include information about the target store, such as which products are worth purchasing. Based on this, the terminal can render the media content in the preview stream. At this point, the user can view the preset media content corresponding to the target store in a preset area near the image of the target store on the phone. After the user has viewed the preset media content corresponding to the target store, he or she can have a general understanding of the target store.
[0624] Different digital maps can be set for different places, so that when the user moves to other places, he can also obtain the preset media content corresponding to the real-time posture based on the method of rendering media content provided in the embodiment of the present disclosure, and render the media content in the preview stream.
[0625] Through the embodiments of the present disclosure, in some scenes with weak textures or high texture similarity (such as in corridors or where the wall occupies a large area of the image to be queried, etc.), candidate reference images that match the image to be queried can be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if the texture in the image to be queried is weak or there are fewer textures, candidate reference images with higher accuracy can still be queried based on the text fields. The initial posture of the terminal obtained by performing posture solution processing based on the candidate reference images with higher accuracy is more accurate. Retrieval and precise positioning based on text field retrieval and feature matching of text area images can utilize text semantic information in the scene, thereby improving the positioning success rate in some areas with similar textures or repeated textures, and because the 2D-3D correspondence of the reference image of the text area image is utilized, the positioning accuracy is made higher.
[0626] Through the embodiments of the present disclosure, text fields can be imperceptibly integrated into visual features, so the recall rate and precision of image retrieval will be higher, and the process is imperceptible to the user, the positioning process is also more intelligent, and the user experience is better.
[0627] Through the embodiments of the present disclosure, even if there are many similar environments at different locations in the target location, and there are many similar images in different images taken at different locations in these locations, it is still possible to query candidate reference images that match the image to be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if there is interference from many similar images, the accuracy of the candidate reference images queried based on the text fields is higher, and the initial pose of the terminal obtained by performing pose solution processing based on the candidate reference images with higher accuracy is more accurate.
[0628] It should be noted that for Figure 19 and Figure 20 In the method for determining the pose shown, since the terminal may not process the N text fields contained in the query image, the terminal may not issue a prompt signal in real time to prompt the user to capture the query image containing text. This means that the query image sent to the server may not contain text. In this case, the server cannot detect the text fields in the query image. In this case, it is impossible to determine candidate reference images based on the text fields. In this case, candidate reference images can be determined according to some methods in the related art. For example, global features can be extracted from the query image and used to determine candidate reference images.
[0629] If the query image does not contain text, the method for determining the subsequent reference image of the query image may include the following scheme:
[0630] In one possible implementation, global image features of each reference image are obtained; global image features of the image to be queried are determined; distances between the global image features of each reference image and the global image features of the image to be queried are determined; and the reference image with the smallest distance is determined as the target candidate reference image.
[0631] In a possible implementation, image similarities between each reference image and the query image are determined; and a candidate reference image with the greatest image similarity is determined as a target candidate reference image.
[0632] In a possible implementation, positioning information sent by a terminal is received; shooting positions corresponding to respective reference images are acquired; and a target reference image whose shooting position matches the positioning information is determined among the reference images.
[0633] Based on the above, if Figure 21 FIG. 1 is a flow chart of a method for determining a posture provided by an embodiment of the present disclosure. The flow of the method for determining a posture may include:
[0634] Step S2101, capture video stream.
[0635] Step S2102: extract the image to be queried from the video stream.
[0636] Step S2103: Perform text detection on the query image.
[0637] Step S2104: determine whether a text box is detected.
[0638] Step S2105: If no text box is detected, image retrieval is performed based on the global features of the image to be queried to obtain the target environment image.
[0639] Step S2106: If a text box is detected, the characters in the text box are recognized.
[0640] Step S2107: judging whether the character corresponding to the query image is correctly recognized by the characters corresponding to the upper and lower video frames of the query image. If not, proceeding to step S2105.
[0641] Step S2108: If the character recognition result corresponding to the query image is correct, image enhancement processing is performed on the text area image.
[0642] Step S2109: extracting image key points of the text area image after image enhancement processing.
[0643] Step S2110: If the character recognition result corresponding to the image to be queried is correct, an image search is performed based on the character recognition result to obtain the target environment image.
[0644] Step S2111: Key point matching is performed based on the image key points of the text region image after image enhancement processing and the image key points of the text region image of the target environment image. Alternatively, key point matching is performed based on the image key points of the target environment image determined in step S2105 and the image key points of the query image.
[0645] Step S2112: establishing a target 2D-3D correspondence relationship based on the key point matching results.
[0646] Step S2113: Perform pose estimation based on the target 2D-3D correspondence.
[0647] The processing of some terminals and the processing of the server in the embodiments disclosed in this disclosure are similar to the processing of some terminals and the processing of the server in the previously disclosed embodiments. For some of the places that can be shared, no excessive description is made in the embodiments disclosed in this disclosure. You can refer to the description of the processing of the terminal and the processing of the server in the previously disclosed embodiments. It should be noted that the processing of each step in S2101-S2113 can be performed by the terminal or by the server, and there are many possible combinations, which are not listed here one by one. In any of the above possible interactive combinations, when implementing the above steps based on the concept of the present invention, those skilled in the art can construct the communication process between the server and the terminal, such as what necessary information is interacted and transmitted between the server and the terminal; this is not exhaustively listed or detailed in the present invention.
[0648] Through the embodiments of the present disclosure, in some scenes with weak textures or high texture similarity (such as in corridors or where the wall occupies a large area of the image to be queried, etc.), candidate reference images that match the image to be queried can be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if the texture in the image to be queried is weak or there are fewer textures, candidate reference images with higher accuracy can still be queried based on the text fields. The initial posture of the terminal obtained by performing posture solution processing based on the candidate reference images with higher accuracy is more accurate. Retrieval and precise positioning based on text field retrieval and feature matching of text area images can utilize text semantic information in the scene, thereby improving the positioning success rate in some areas with similar textures or repeated textures, and because the 2D-3D correspondence of the reference image of the text area image is utilized, the positioning accuracy is made higher.
[0649] Through the embodiments of the present disclosure, text fields can be imperceptibly integrated into visual features, so the recall rate and precision of image retrieval will be higher, and the process is imperceptible to the user, the positioning process is also more intelligent, and the user experience is better.
[0650] Through the embodiments of the present disclosure, even if there are many similar environments at different locations in the target location, and there are many similar images in different images taken at different locations in these locations, it is still possible to query candidate reference images that match the image to be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if there is interference from many similar images, the accuracy of the candidate reference images queried based on the text fields is higher, and the initial pose of the terminal obtained by performing pose solution processing based on the candidate reference images with higher accuracy is more accurate.
[0651] In summary, the methods for determining posture provided by the embodiments of the present application can be divided into two categories. The difference between these two categories is that in one category, the terminal determines the N text fields contained in the query image and sends the query image and text fields to the server, while in the other category, the terminal sends the query image to the server, and the server determines the N text fields contained in the query image. The following is an integrated description of these two methods for determining posture, which can be described as follows:
[0652] like Figure 22 FIG. 1 is a flow chart of a method for determining a posture provided by an embodiment of the present disclosure. The flow of the method for determining a posture may include:
[0653] Step S2201: Acquire a query image at a first location.
[0654] Step S2202: Determine N text fields contained in the image to be queried.
[0655] Step S2203 : Based on the pre-stored correspondence between reference images and text fields, candidate reference images are determined according to the N text fields.
[0656] Step S2204: Determine the initial posture of the terminal at the first position based on the image to be queried and the candidate reference images.
[0657] It should be noted that step S2202 can be executed by the terminal or by the server.
[0658] When step S2202 is executed by the terminal, after determining N text fields, the N text fields and the image to be queried need to be sent to the server. For specific processing, please refer to the relevant content in steps S1702-S1703, S1603 and S802.
[0659] When step S2203 is executed by the server, the terminal sends the image to be queried to the server, and then the server processes the image to be queried to determine N text fields. For specific processing, please refer to the relevant content in step S2002.
[0660] An exemplary embodiment of the present disclosure provides an offline calibration method, which can be performed before the actual online positioning process to determine some correspondences required for the online positioning process. The method can be applied to a server, such as Figure 23 As shown, the processing flow of the method may include the following steps:
[0661] Step S2301: Acquire a pre-collected reference image.
[0662] Step S2302: Determine the text fields contained in each reference image.
[0663] Step S2303: Store the text fields and the reference images in correspondence.
[0664] The above processing steps have been described in the corresponding processing steps of the online positioning process described in the previous embodiment, and will not be repeated here. Please refer to the content described in the previous embodiment.
[0665] An exemplary embodiment of the present disclosure further provides an offline calibration method, such as Figure 24 As shown, the processing flow of the method may include the following steps:
[0666] Step S2401 : Acquire a plurality of pre-collected reference images, perform text detection processing on each reference image, and determine the text region image contained in each reference image.
[0667] The multiple reference images collected in advance can be environment images taken at various locations in the target location. Each time the location is changed, an offline calibration process can be performed.
[0668] Step S2402: Identify the text fields contained in each text region image.
[0669] Step S2403: Create a search index based on each text field and register the search index into the database.
[0670] Step S2404: performing image enhancement processing on each text region image.
[0671] Step S2405: extracting image key points of each text region image after image enhancement processing.
[0672] Step S2406 , obtaining a 3D point cloud, calibrating the correspondence between the extracted image key points and the 3D points in the 3D point cloud, and registering the calibrated correspondence into the database.
[0673] Optionally, the method provided by the embodiment of the present disclosure may also include: determining the 2D-3D correspondence of each text area image based on the 2D points of the text area image contained in each reference image and the 2D-3D correspondence of each reference image acquired in advance; and storing the 2D-3D correspondence of each text area image.
[0674] During implementation, the server can obtain pre-established 2D-3D correspondences for each reference image. The 2D-3D correspondences include a large number of 3D points and corresponding 2D points in the reference images. Each 3D point corresponds to a physical point in the target location, and each 3D point has three-dimensional position information of a corresponding physical point in real space. After determining the text region image from each reference image, the 2D points of each text region image can be obtained. Based on the 2D-3D correspondences of each text region image and the reference image to which each text region image belongs, the 2D-3D correspondences of each text region image can be determined. After the 2D-3D correspondences of each text region image are established, they can be stored in the server's database.
[0675] The server may pre-establish three-dimensional position information of the physical object point corresponding to each pixel point in the reference image in the actual space, which is recorded as a 2D-3D correspondence relationship.
[0676] Based on an automatic calibration method for OCR regions (text image regions) and 3D point clouds (2D-3D correspondences between reference images), this method can automatically calibrate and align text box regions in the image library with the 2D-3D correspondences of existing reference images offline, providing a data foundation for subsequent precise positioning. This offline calibration method provides a solid data foundation for the subsequent online positioning process.
[0677] Through the embodiments of the present disclosure, even if there are many similar environments at different locations in the target location, and there are many similar images in different images taken at different locations in these locations, it is still possible to query candidate reference images that match the image to be queried based on the text fields contained in the image to be queried and the text fields contained in different reference images. Even if there is interference from many similar images, the accuracy of the candidate reference images queried based on the text fields is higher, and the initial pose of the terminal obtained by performing pose solution processing based on the candidate reference images with higher accuracy is more accurate.
[0678] Another exemplary embodiment of the present disclosure provides a device for determining a posture, such as Figure 25 As shown, the device includes:
[0679] The acquisition module 2501 is used to acquire the image to be queried at a first position, wherein the scene at the first position includes the scene in the image to be queried; there is text in the image to be queried; specifically, it can implement the acquisition function in the above step S601 and other implicit steps.
[0680] The determination module 2502 is used to determine N text fields contained in the image to be queried, where N is greater than or equal to 1. Specifically, it can implement the determination function in the above step S602 and other implicit steps.
[0681] The sending module 2503 is used to send the N text fields and the image to be queried to the server; specifically, it can implement the sending function in the above step S603 and other implicit steps.
[0682] The receiving module 2504 is used to receive the initial posture of the terminal at the first position returned by the server; wherein, the initial posture is determined by the server based on the N text fields and the image to be queried; specifically, it can implement the receiving function in the above-mentioned step S604, as well as other implicit steps.
[0683] In a possible implementation, the obtaining module 2501 is configured to:
[0684] capturing a first initial image;
[0685] When no text exists in the first initial image, displaying or announcing a first prompt message; wherein the first prompt message is used to indicate that no text is detected in the first initial image and prompt the user to move the position of the terminal or adjust the shooting angle of the terminal;
[0686] When a second initial image containing text is captured at the first position, the second initial image is determined as the image to be queried.
[0687] In a possible implementation, the obtaining module 2501 is configured to:
[0688] capturing a third initial image;
[0689] Determine a text area image contained in the third initial image by performing text detection processing on the third initial image;
[0690] When the text area image included in the third initial image does not meet the preferred image condition, displaying or announcing a second prompt message; wherein the second prompt message is used to indicate that the text area image included in the third initial image does not meet the preferred image condition and prompting the user to move the terminal in the direction where the physical text is located;
[0691] until a fourth initial image containing a text region image that meets the preferred image condition is captured at the first position, determining the fourth initial image as the image to be queried;
[0692] The preferred image conditions include one or more of the following conditions:
[0693] The size of the text area image is greater than or equal to a size threshold;
[0694] The clarity of the text area image is greater than or equal to a clarity threshold;
[0695] The texture complexity of the text region image is less than or equal to a complexity threshold.
[0696] In a possible implementation, the obtaining module 2501 is configured to:
[0697] capturing a fifth initial image;
[0698] Determining N text fields included in the fifth initial image;
[0699] Acquire M text fields contained in a reference query image, wherein a time interval between a capture time of the reference query image and a capture time of the fifth initial image is less than a time threshold, and M is greater than or equal to 1;
[0700] When any text field included in the fifth initial image is inconsistent with each of the M text fields, a third prompt message is displayed or announced by voice; wherein the third prompt message is used to indicate that an incorrect text field is recognized in the fifth initial image and prompt the user to move the position of the terminal or adjust the shooting angle of the terminal;
[0701] When each text field contained in the sixth initial image captured at the first position belongs to the M text fields, the sixth initial image is determined as the image to be queried.
[0702] In a possible implementation, the obtaining module 2501 is configured to:
[0703] Capturing a first image of a current scene at a first position; the first image includes text;
[0704] Performing text detection processing on the first image to obtain at least one text area image;
[0705] The at least one text region image contained in the first image is used as a query image.
[0706] In a possible implementation, the determining module 2502 is further configured to determine a location of the text region image in the query image;
[0707] The sending module 2503 is also used to send the location area to the server; the initial pose is determined by the server based on the N text fields and the image to be queried, including: the initial pose is determined by the server based on the location area of the text area image in the image to be queried, the N text fields and the image to be queried.
[0708] In a possible implementation, the acquisition module 2501 is further configured to acquire the location information of the terminal;
[0709] The sending module 2503 is also used to send the positioning information to the server; the initial posture is determined by the server based on the N text fields and the image to be queried, including: the initial posture is determined by the server based on the N text fields, the image to be queried and the positioning information.
[0710] In a possible implementation, the acquisition module 2501 is further configured to acquire a posture change of the terminal;
[0711] The determination module 2502 is further configured to determine a real-time posture according to the initial posture and the posture change of the terminal.
[0712] In a possible implementation, the acquisition module 2501 is further configured to acquire a preview stream of the current scene;
[0713] The determining module 2502 is further configured to determine, based on the real-time posture, preset media content contained in the digital map corresponding to the scene in the preview stream;
[0714] The device further comprises:
[0715] A rendering module is configured to render the media content in the preview stream.
[0716] It should be noted that the above-mentioned acquisition module 2501, determination module 2502, sending module 2503, and receiving module 2504 can be implemented by a processor, or by a processor in conjunction with a transceiver.
[0717] Another exemplary embodiment of the present disclosure provides a device for determining a posture, such as Figure 26 As shown, the device includes:
[0718] Receiving module 2601 is used to receive the image to be queried and N text fields contained in the image to be queried sent by the terminal, where N is greater than or equal to 1; the image to be queried is obtained based on the image captured by the terminal at the first position; the scene at the first position includes the scene in the image to be queried; specifically, it can implement the receiving function in the above-mentioned step S701 and other implicit steps.
[0719] The determination module 2602 is used to determine candidate reference images according to the N text fields based on the pre-stored correspondence between reference images and text fields; specifically, it can implement the determination function in the above step S702 and other implicit steps.
[0720] The determination module 2602 is used to determine the initial posture of the terminal at the first position based on the query image and the candidate reference image; specifically, it can implement the determination function in the above step S703 and other implicit steps.
[0721] The sending module 2603 is used to send the determined initial posture to the terminal; specifically, it can implement the sending function in the above step S704 and other implicit steps.
[0722] In a possible implementation, the determining module 2602 is configured to:
[0723] Determining a target reference image from the candidate reference images; wherein the scene at the first position includes the scene in the target reference image;
[0724] An initial posture of the terminal at the first position is determined according to the image to be queried and the target reference image.
[0725] In a possible implementation, the determining module 2602 is configured to:
[0726] Determining a 2D-2D correspondence between the query image and the target reference image;
[0727] An initial posture of the terminal at the first position is determined according to the 2D-2D correspondence and a preset 2D-3D correspondence of the target reference image.
[0728] In a possible implementation, the receiving module 2601 is further configured to receive a location area of the text area image sent by the terminal in the image to be queried;
[0729] The determining module 2602 is configured to determine a target text area image contained in the image to be queried based on the location area;
[0730] The device further comprises:
[0731] An acquisition module, configured to acquire a text region image contained in the target reference image;
[0732] The determining module 2602 is configured to determine a 2D-2D correspondence between the target text region image and the text region image included in the target reference image.
[0733] In a possible implementation, the determining module 2602 is configured to:
[0734] Determining the image similarity between each candidate reference image and the query image;
[0735] Among the candidate reference images, a first preset number of target reference images having the highest image similarity are determined.
[0736] In a possible implementation, the determining module 2602 is configured to:
[0737] Obtaining global image features of each candidate reference image;
[0738] Determining global image features of the image to be queried;
[0739] Determining the distances between the global image features of each candidate reference image and the global image features of the query image;
[0740] A second preset number of target reference images with the smallest distances are determined among the candidate reference images.
[0741] In a possible implementation, the determining module 2602 is configured to:
[0742] receiving positioning information sent by the terminal;
[0743] Obtaining the shooting positions corresponding to each candidate reference image;
[0744] Among the candidate reference images, a target reference image whose shooting position matches the positioning information is determined.
[0745] In a possible implementation, when N is greater than 1, the determining module 2602 is configured to:
[0746] Among the candidate reference images, a target reference image including the N text fields is determined.
[0747] In a possible implementation, the determining module 2602 is configured to:
[0748] When the number of the candidate reference images is equal to 1, the candidate reference image is determined as the target reference image.
[0749] In a possible implementation, the determining module 2602 is configured to:
[0750] Inputting N text fields contained in the image to be queried into a pre-trained text classifier to obtain a text type of each text field contained in the image to be queried;
[0751] Determine that the text type is a text field of a preset significant type;
[0752] Based on the pre-stored correspondence between reference images and text fields, a candidate reference image corresponding to the salient type of text field is searched.
[0753] It should be noted that the above-mentioned receiving module 2601, determining module 2602, and sending module 2603 can be implemented by a processor, or by a processor in conjunction with a memory and a transceiver.
[0754] Regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0755] Another exemplary embodiment of the present disclosure provides a device for determining a posture, such as Figure 27 As shown, the device includes:
[0756] The acquisition module 2701 is used to acquire the image to be queried at a first position, wherein the scene at the first position includes the scene in the image to be queried; specifically, it can implement the acquisition function in the above step S1901 and other implicit steps.
[0757] The sending module 2702 is used to send the image to be queried to the server so that the server can determine the N text fields included in the image to be queried, and determine the initial posture of the terminal at the first position based on the N text fields and the image to be queried, where N is greater than or equal to 1; specifically, it can implement the sending function in the above-mentioned step S1902, as well as other implicit steps.
[0758] The receiving module 2703 is used to receive the initial position of the terminal at the first position returned by the server; specifically, it can implement the receiving function in the above step S1903 and other implicit steps.
[0759] In a possible implementation, the acquisition module 2701 is further configured to acquire the location information of the terminal;
[0760] The sending module 2702 is also used to send positioning information to the server; determining the initial posture of the terminal at the first position based on the N text fields and the image to be queried includes: determining the initial posture of the terminal at the first position based on the N text fields, the image to be queried and the positioning information.
[0761] In a possible implementation, the acquisition module 2701 is further configured to acquire a posture change of the terminal;
[0762] The device further includes a determination module for determining a real-time posture according to the initial posture and the posture change of the terminal.
[0763] In a possible implementation, the acquisition module 2701 is further configured to acquire a preview stream of the current scene;
[0764] The determining module is further configured to determine, based on the real-time posture, preset media content contained in the digital map corresponding to the scene in the preview stream;
[0765] The device further comprises:
[0766] A rendering module is configured to render the media content in the preview stream.
[0767] It should be noted that the above-mentioned acquisition module 2701, sending module 2702, and receiving module 2703 can be implemented by a processor, or by a processor in conjunction with a transceiver.
[0768] Another exemplary embodiment of the present disclosure provides a device for determining a posture, such as Figure 28 As shown, the device includes:
[0769] Receiving module 2801 is used to receive the image to be queried sent by the terminal, wherein the image to be queried is obtained based on the image captured by the terminal at the first posit...
Claims
1. A method for determining a posture, characterized in that: The method comprises: The terminal obtains an image to be queried at a first position, wherein the scene at the first position includes the scene in the image to be queried; and there is text in the image to be queried; Determining N text fields contained in the image to be queried, where N is greater than or equal to 1; Sending the N text fields and the image to be queried to a server; Receiving an initial position of the terminal at the first position returned by the server; Among them, the server determines the target text area image contained in the image to be queried, and the N text fields are located in the target text area image; based on the pre-stored correspondence between the reference image and the text field, determines the candidate reference image according to the N text fields; determines the target reference image in the candidate reference image; obtains the text area image contained in the target reference image; determines the 2D-2D correspondence between the target text area image and the text area image contained in the target reference image; based on the 2D-2D correspondence and the 2D-3D correspondence of the text area image included in the target reference image, determines the initial posture of the terminal at the first position.
2. The method according to claim 1, characterized in that The terminal obtains an image to be queried at a first location, including: capturing a first initial image; When no text exists in the first initial image, displaying or announcing a first prompt message; wherein the first prompt message is used to indicate that no text is detected in the first initial image and prompt the user to move the position of the terminal or adjust the shooting angle of the terminal; When the terminal captures a second initial image containing text at the first position, the second initial image is determined as the image to be queried.
3. The method according to claim 1, characterized in that The terminal obtains an image to be queried at a first location, including: capturing a third initial image; Determine a text area image contained in the third initial image by performing text detection processing on the third initial image; When the text area image included in the third initial image does not meet the preferred image condition, displaying or announcing a second prompt message; wherein the second prompt message is used to indicate that the text area image included in the third initial image does not meet the preferred image condition and prompting the user to move the terminal in the direction where the physical text is located; until the terminal captures a fourth initial image at the first position, the image of the text region containing the image meeting the preferred image condition, and determines the fourth initial image as the image to be queried; The preferred image conditions include one or more of the following conditions: The size of the text area image is greater than or equal to a size threshold; The clarity of the text area image is greater than or equal to a clarity threshold; The texture complexity of the text region image is less than or equal to a complexity threshold.
4. The method according to claim 1, wherein The terminal obtains an image to be queried at a first location, including: capturing a fifth initial image; Determining N text fields included in the fifth initial image; Acquire M text fields contained in a reference query image, wherein a time interval between a capture time of the reference query image and a capture time of the fifth initial image is less than a time threshold, and M is greater than or equal to 1; When any text field included in the fifth initial image is inconsistent with each of the M text fields, a third prompt message is displayed or announced by voice; wherein the third prompt message is used to indicate that an incorrect text field is recognized in the fifth initial image and prompt the user to move the position of the terminal or adjust the shooting angle of the terminal; When each text field contained in the sixth initial image photographed by the terminal at the first position belongs to the M text fields, the sixth initial image is determined as the image to be queried.
5. The method according to any one of claims 1 to 4, characterized in that The terminal obtains an image to be queried at a first location, including: The terminal captures a first image of the current scene at the first position; the first image contains text; Performing text detection processing on the first image to obtain at least one text area image; The at least one text region image contained in the first image is used as a query image.
6. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Determine a location area of the text area image in the image to be queried; The location area is sent to the server.
7. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Obtaining location information of the terminal; sending the positioning information to the server; The server further determines the initial posture based on the positioning information.
8. The method according to any one of claims 1 to 4, characterized in that After receiving the initial posture returned by the server, the method further includes: Obtaining a posture change of the terminal; Determine a real-time posture according to the initial posture and the posture change of the terminal.
9. The method according to claim 8, characterized in that After receiving the initial posture returned by the server, the method further includes: Get the preview stream of the current scene; Determining, based on the real-time posture, preset media content contained in a digital map corresponding to the scene in the preview stream; The media content is rendered in the preview stream.
10. A method for determining a posture, characterized in that: The method comprises: Receiving an image to be queried and N text fields contained in the image to be queried, sent by a terminal, where N is greater than or equal to 1; the image to be queried is obtained based on an image captured by the terminal at a first position; and the scene at the first position includes the scene in the image to be queried; Determining a target text area image contained in the image to be queried, wherein the N text fields are located in the target text area image; Based on a pre-stored correspondence between reference images and text fields, determining candidate reference images according to the N text fields, determining a target reference image among the candidate reference images, and acquiring a text region image contained in the target reference image; determining a 2D-2D correspondence between the target text region image and the text region image included in the target reference image; determining an initial posture of the terminal at the first position according to the 2D-2D correspondence and a 2D-3D correspondence of the text area image contained in the target reference image; The initial posture is sent to the terminal.
11. The method according to claim 10, characterized in that The method further comprises: receiving a location area of the target text area image in the image to be queried, sent by the terminal; The determining of the target text area image contained in the query image includes: Based on the location area, the target text area image contained in the image to be queried is determined.
12. The method according to claim 10, characterized in that The determining of a target reference image from the candidate reference images comprises: Determining the image similarity between each candidate reference image and the query image; The candidate reference images whose image similarity is greater than or equal to a preset similarity threshold are determined as target reference images.
13. The method according to claim 10, characterized in that The determining of a target reference image from the candidate reference images comprises: Obtaining global image features of each candidate reference image; Determining global image features of the image to be queried; Determining the distances between the global image features of each candidate reference image and the global image features of the query image; The candidate reference image whose distance is less than or equal to the preset distance threshold is determined as the target reference image.
14. The method according to claim 10, characterized in that The determining of a target reference image from the candidate reference images comprises: receiving positioning information sent by the terminal; Obtaining the shooting positions corresponding to each candidate reference image; Among the candidate reference images, a target reference image whose shooting position matches the positioning information is determined.
15. The method according to claim 10, characterized in that When N is greater than 1, determining a target reference image from the candidate reference images includes: Among the candidate reference images, a target reference image including the N text fields is determined.
16. The method according to claim 10, characterized in that The determining of a target reference image from the candidate reference images comprises: When the number of the candidate reference images is equal to 1, the candidate reference image is determined as the target reference image.
17. The method according to any one of claims 10 to 16, characterized in that The determining of candidate reference images according to the N text fields based on the pre-stored correspondence between reference images and text fields includes: Inputting N text fields contained in the image to be queried into a pre-trained text classifier to obtain a text type of each text field contained in the image to be queried; Determine that the text type is a text field of a preset significant type; Based on the pre-stored correspondence between reference images and text fields, a candidate reference image corresponding to the salient type of text field is searched.
18. A device for determining a posture, characterized in that: The device comprises: An acquisition module is configured to acquire an image to be queried at a first position, wherein the scene at the first position includes the scene in the image to be queried; and the image to be queried contains text; a determination module, configured to determine N text fields contained in the image to be queried, where N is greater than or equal to 1; A sending module, configured to send the N text fields and the image to be queried to a server; A receiving module, configured to receive an initial position of the terminal at the first position returned by the server; Among them, the server determines the target text area image contained in the image to be queried, and the N text fields are located in the target text area image; based on the pre-stored correspondence between the reference image and the text field, determines the candidate reference image according to the N text fields; determines the target reference image in the candidate reference image; obtains the text area image contained in the target reference image; determines the 2D-2D correspondence between the target text area image and the text area image contained in the target reference image; based on the 2D-2D correspondence and the 2D-3D correspondence of the text area image included in the target reference image, determines the initial posture of the terminal at the first position.
19. The device according to claim 18, characterized in that The acquisition module is used to: capturing a first initial image; When no text exists in the first initial image, displaying or announcing a first prompt message; wherein the first prompt message is used to indicate that no text is detected in the first initial image and prompt the user to move the position of the terminal or adjust the shooting angle of the terminal; When a second initial image containing text is captured at the first position, the second initial image is determined as the image to be queried.
20. The device according to claim 18, characterized in that The acquisition module is used to: capturing a third initial image; Determine a text area image contained in the third initial image by performing text detection processing on the third initial image; When the text area image included in the third initial image does not meet the preferred image condition, displaying or announcing a second prompt message; wherein the second prompt message is used to indicate that the text area image included in the third initial image does not meet the preferred image condition and prompting the user to move the terminal in the direction where the physical text is located; until a fourth initial image containing a text region image that meets the preferred image condition is captured at the first position, determining the fourth initial image as the image to be queried; The preferred image conditions include one or more of the following conditions: The size of the text area image is greater than or equal to a size threshold; The clarity of the text area image is greater than or equal to a clarity threshold; The texture complexity of the text region image is less than or equal to a complexity threshold.
21. The device according to claim 18, characterized in that The acquisition module is used to: capturing a fifth initial image; Determining N text fields included in the fifth initial image; Acquire M text fields contained in a reference query image, wherein a time interval between a capture time of the reference query image and a capture time of the fifth initial image is less than a time threshold, and M is greater than or equal to 1; When any text field included in the fifth initial image is inconsistent with each of the M text fields, a third prompt message is displayed or announced by voice; wherein the third prompt message is used to indicate that an incorrect text field is recognized in the fifth initial image and prompt the user to move the position of the terminal or adjust the shooting angle of the terminal; When each text field contained in the sixth initial image captured at the first position belongs to the M text fields, the sixth initial image is determined as the image to be queried.
22. The device according to any one of claims 18 to 21, characterized in that The acquisition module is used to: Capturing a first image of the current scene at the first position; the first image containing text; Performing text detection processing on the first image to obtain at least one text area image; The at least one text region image contained in the first image is used as a query image.
23. The device according to any one of claims 18 to 21, characterized in that The determination module is further configured to determine a location area of the text area image in the image to be queried; The sending module is further configured to send the location area to the server.
24. The device according to any one of claims 18 to 21, characterized in that The acquisition module is further configured to acquire the location information of the terminal; The sending module is further configured to send the positioning information to the server; The server further determines the initial posture based on the positioning information.
25. The device according to any one of claims 18 to 21, characterized in that The acquisition module is further configured to acquire the posture change of the terminal; The determination module is further configured to determine a real-time posture according to the initial posture and the posture change of the terminal.
26. The device according to claim 25, characterized in that The acquisition module is further used to obtain a preview stream of the current scene; The determining module is further configured to determine, based on the real-time posture, preset media content contained in the digital map corresponding to the scene in the preview stream; The device further comprises: A rendering module is configured to render the media content in the preview stream.
27. A device for determining posture, characterized in that The device comprises: a receiving module, configured to receive an image to be queried and N text fields contained in the image to be queried, sent by a terminal, where N is greater than or equal to 1; the image to be queried is obtained based on an image captured by the terminal at a first position; and the scene at the first position includes the scene in the image to be queried; a determination module, configured to determine a target text region image contained in the image to be queried, wherein the N text fields are located in the target text region image; The determining module is further configured to determine candidate reference images according to the N text fields based on a pre-stored correspondence between reference images and text fields; and determine a target reference image from the candidate reference images; The device further includes an acquisition module, which is used to acquire a text area image contained in the target reference image; The determination module is further configured to determine a 2D-2D correspondence between the target text area image and the text area image included in the target reference image; and determine an initial posture of the terminal at the first position based on the 2D-2D correspondence and the 2D-3D correspondence of the text area image included in the target reference image; A sending module is used to send the initial posture to the terminal.
28. The device according to claim 27, characterized in that The receiving module is further configured to receive a location area of the target text area image in the image to be queried, which is sent by the terminal; The determining module is configured to determine the target text area image contained in the image to be queried based on the location area.
29. The device according to claim 27, characterized in that The determining module is configured to: Determining the image similarity between each candidate reference image and the query image; Among the candidate reference images, a first preset number of target reference images having the highest image similarity are determined.
30. The device according to claim 27, characterized in that The determining module is configured to: Obtaining global image features of each candidate reference image; Determining global image features of the image to be queried; Determining the distances between the global image features of each candidate reference image and the global image features of the query image; A second preset number of target reference images with the smallest distances are determined among the candidate reference images.
31. The device according to claim 27, characterized in that The determining module is configured to: receiving positioning information sent by the terminal; Obtaining the shooting positions corresponding to each candidate reference image; Among the candidate reference images, a target reference image whose shooting position matches the positioning information is determined.
32. The device according to claim 27, wherein When N is greater than 1, the determining module is configured to: Among the candidate reference images, a target reference image including the N text fields is determined.
33. The device according to claim 27, characterized in that The determining module is configured to: When the number of the candidate reference images is equal to 1, the candidate reference image is determined as the target reference image.
34. The device according to any one of claims 27 to 33, characterized in that The determining module is configured to: Inputting N text fields contained in the image to be queried into a pre-trained text classifier to obtain a text type of each text field contained in the image to be queried; Determine that the text type is a text field of a preset significant type; Based on the pre-stored correspondence between reference images and text fields, a candidate reference image corresponding to the salient type of text field is searched.
35. A system for determining posture, characterized in that The system includes a terminal and a server, wherein: The terminal is configured to obtain an image to be queried at a first location, wherein a scene at the first location includes a scene in the image to be queried; the image to be queried contains text; determine N text fields contained in the image to be queried, wherein N is greater than or equal to 1; send the N text fields and the image to be queried to the server; and receive an initial pose returned by the server; The server is configured to receive an image to be queried and N text fields contained in the image to be queried sent by the terminal; determine a target text area image contained in the image to be queried, wherein the N text fields are located in the target text area image; determine a candidate reference image according to the N text fields based on a pre-stored correspondence between reference images and text fields, determine a target reference image among the candidate reference images, and obtain the text area image contained in the target reference image; determine a 2D-2D correspondence between the target text area image and the text area image contained in the target reference image; determine an initial posture of the terminal at the first position based on the 2D-2D correspondence and the 2D-3D correspondence of the text area image contained in the target reference image; and send the initial posture to the terminal.
36. A method for determining a posture, characterized in that The method comprises: The terminal obtains an image to be queried at a first position, wherein the scene at the first position includes the scene in the image to be queried; Sending the image to be queried to the server; receiving an initial position of the terminal at the first position returned by the server; In which, the server determines N text fields contained in the image to be queried, N is greater than or equal to 1; based on the pre-stored correspondence between reference images and text fields, determines a candidate reference image according to the N text fields; determines a target reference image in the candidate reference image; obtains a text area image contained in the target reference image; determines a target text area image contained in the image to be queried, wherein the N text fields are located in the target text area image; determines a 2D-2D correspondence between the target text area image and the text area image contained in the target reference image; and determines an initial posture of the terminal at the first position based on the 2D-2D correspondence and the 2D-3D correspondence of the text area image included in the target reference image.
37. The method according to claim 36, wherein The method further comprises: Obtaining location information of the terminal; sending the positioning information to the server; The server is further configured to determine an initial position of the terminal at the first location based on the positioning information.
38. The method according to claim 36 or 37, characterized in that After receiving the initial posture returned by the server, the method further includes: Obtaining a posture change of the terminal; Determine a real-time posture according to the initial posture and the posture change of the terminal.
39. The method according to claim 38, wherein After receiving the initial posture returned by the server, the method further includes: Get the preview stream of the current scene; Determining, based on the real-time posture, preset media content contained in a digital map corresponding to the scene in the preview stream; The media content is rendered in the preview stream.
40. A method for determining a posture, characterized in that: The method comprises: Receiving an image to be queried sent by a terminal, wherein the image to be queried is obtained based on an image captured by the terminal at a first location; and a scene at the first location includes a scene in the image to be queried; Determining N text fields contained in the image to be queried, where N is greater than or equal to 1; Determining a target text area image contained in the image to be queried, wherein the N text fields are located in the target text area image; Based on a pre-stored correspondence between reference images and text fields, determining candidate reference images according to the N text fields, determining a target reference image among the candidate reference images, and acquiring a text region image contained in the target reference image; determining a 2D-2D correspondence between the target text region image and the text region image included in the target reference image; determining an initial posture of the terminal at the first position according to the 2D-2D correspondence and a 2D-3D correspondence of the text area image included in the target reference image; The initial posture is sent to the terminal.
41. The method according to claim 40, characterized in that The determining of a target reference image from the candidate reference images comprises: Determining the image similarity between each candidate reference image and the query image; The candidate reference images whose image similarity is greater than or equal to a preset similarity threshold are determined as target reference images.
42. The method according to claim 40, wherein The determining of a target reference image from the candidate reference images comprises: Obtaining global image features of each candidate reference image; Determining global image features of the image to be queried; Determining the distances between the global image features of each candidate reference image and the global image features of the query image; The candidate reference image whose distance is less than or equal to the preset distance threshold is determined as the target reference image.
43. The method according to claim 40, wherein The determining of a target reference image from the candidate reference images comprises: receiving positioning information sent by the terminal; Obtaining the shooting positions corresponding to each candidate reference image; Among the candidate reference images, a target reference image whose shooting position matches the positioning information is determined.
44. The method according to claim 40, wherein When N is greater than 1, determining a target reference image from the candidate reference images includes: Among the candidate reference images, a target reference image including the N text fields is determined.
45. The method according to claim 40, wherein The determining of a target reference image from the candidate reference images comprises: When the number of the candidate reference images is equal to 1, the candidate reference image is determined as the target reference image.
46. The method according to any one of claims 40 to 45, characterized in that The determining of candidate reference images according to the N text fields based on the pre-stored correspondence between reference images and text fields includes: Inputting N text fields contained in the image to be queried into a pre-trained text classifier to obtain a text type of each text field contained in the image to be queried; Determine that the text type is a text field of a preset significant type; Based on the pre-stored correspondence between reference images and text fields, a candidate reference image corresponding to the salient type of text field is searched.
47. A device for determining posture, characterized in that The device comprises: An acquisition module, configured to acquire an image to be queried at a first location, wherein a scene at the first location includes a scene in the image to be queried, and the image to be queried is taken by a terminal; A sending module, configured to send the image to be queried to a server; a receiving module, configured to receive an initial position of the terminal at the first position returned by the server; In which, the server determines N text fields contained in the image to be queried, N is greater than or equal to 1; based on the pre-stored correspondence between reference images and text fields, determines a candidate reference image according to the N text fields; determines a target reference image in the candidate reference image; obtains a text area image contained in the target reference image; determines a target text area image contained in the image to be queried, wherein the N text fields are located in the target text area image; determines a 2D-2D correspondence between the target text area image and the text area image contained in the target reference image; and determines an initial posture of the terminal at the first position based on the 2D-2D correspondence and the 2D-3D correspondence of the text area image included in the target reference image.
48. The device according to claim 47, characterized in that The acquisition module is further configured to acquire the location information of the terminal; The sending module is further configured to send positioning information to the server; The server is further configured to determine an initial position of the terminal at the first location based on the positioning information.
49. The device according to claim 47 or 48, characterized in that The acquisition module is further configured to acquire the posture change of the terminal; The device further includes a determination module for determining a real-time posture according to the initial posture and the posture change of the terminal.
50. The device according to claim 49, characterized in that The acquisition module is further used to obtain a preview stream of the current scene; The determining module is further configured to determine, based on the real-time posture, preset media content contained in the digital map corresponding to the scene in the preview stream; The device further comprises: A rendering module is configured to render the media content in the preview stream.
51. A device for determining posture, characterized in that The device comprises: a receiving module, configured to receive an image to be queried sent by a terminal, wherein the image to be queried is obtained based on an image captured by the terminal at a first position; and a scene at the first position includes a scene in the image to be queried; a determination module, configured to determine N text fields contained in the image to be queried, where N is greater than or equal to 1; The determining module is further configured to determine a target text region image contained in the image to be queried, wherein the N text fields are located in the target text region image; The determination module is further configured to determine candidate reference images according to the N text fields based on a pre-stored correspondence between reference images and text fields, and determine a target reference image among the candidate reference images; The device further includes an acquisition module, which is used to acquire a text area image contained in the target reference image; The determining module is further configured to determine a 2D-2D correspondence between the target text area image and the text area image included in the target reference image; and determine an initial posture of the terminal at the first position based on the 2D-2D correspondence and the 2D-3D correspondence of the text area image included in the target reference image; A sending module is used to send the initial posture to the terminal.
52. The device according to claim 51, characterized in that The determining module is configured to: Determining the image similarity between each candidate reference image and the query image; Among the candidate reference images, a first preset number of target reference images having the highest image similarity are determined.
53. The device according to claim 51, characterized in that The determining module is configured to: Obtaining global image features of each candidate reference image; Determining global image features of the image to be queried; Determining the distances between the global image features of each candidate reference image and the global image features of the query image; A second preset number of target reference images with the smallest distances are determined among the candidate reference images.
54. The device according to claim 51, characterized in that The determining module is configured to: receiving positioning information sent by the terminal; Obtaining the shooting positions corresponding to each candidate reference image; Among the candidate reference images, a target reference image whose shooting position matches the positioning information is determined.
55. The device according to claim 51, characterized in that When N is greater than 1, the determining module is configured to: Among the candidate reference images, a target reference image including the N text fields is determined.
56. The device according to claim 51, characterized in that The determining module is configured to: When the number of the candidate reference images is equal to 1, the candidate reference image is determined as the target reference image.
57. The device according to any one of claims 51 to 56, characterized in that The determining module is configured to: Inputting N text fields contained in the image to be queried into a pre-trained text classifier to obtain a text type of each text field contained in the image to be queried; Determine that the text type is a text field of a preset significant type; Based on the pre-stored correspondence between reference images and text fields, a candidate reference image corresponding to the salient type of text field is searched.
58. A system for determining posture, characterized in that The system includes a terminal and a server, wherein: The terminal is configured to obtain an image to be queried at a first location, wherein a scene at the first location includes a scene in the image to be queried; send the image to be queried to a server; and receive an initial position of the terminal at the first location returned by the server; The server is configured to receive the image to be queried sent by the terminal; determine N text fields contained in the image to be queried, wherein N is greater than or equal to 1; determine a target text area image contained in the image to be queried, wherein the N text fields are located in the target text area image; determine a candidate reference image according to the N text fields based on a pre-stored correspondence between reference images and text fields; determine a target reference image in the candidate reference images; obtain the text area image contained in the target reference image; determine a 2D-2D correspondence between the target text area image and the text area image contained in the target reference image; determine an initial posture of the terminal at the first position based on the 2D-2D correspondence and the 2D-3D correspondence of the text area image included in the target reference image; and send the initial posture to the terminal.
59. A terminal, characterized in that: The terminal includes a processor, a memory, a transceiver, a camera, and a bus, wherein: The processor, the memory, the transceiver and the camera are connected via the bus; The camera is used to capture images; The transceiver is used to receive and send data; The memory is used to store computer programs; The processor is used to control the memory, transceiver and camera, and execute the program stored in the memory to implement the method steps described in any one of claims 1-9 and 36-39.
60. A server, characterized in that The server includes a processor, a memory, a transceiver, and a bus, wherein: The processor, the memory and the transceiver are connected via the bus; The transceiver is used to receive and send data; The memory is used to store computer programs; The processor is used to execute the program stored in the memory to implement the method steps described in any one of claims 10-17 and 40-46.
61. A computer-readable storage medium, characterized in that The computer-readable storage medium comprises instructions, which, when the computer-readable storage medium is executed on a terminal, causes the terminal to execute the method according to any one of claims 1 to 9 and 36 to 39.
62. A computer program product comprising instructions, characterized in that When the computer program product is run on a terminal, the terminal is caused to execute the method according to any one of claims 1 to 9 and 36 to 39.
63. A computer-readable storage medium, characterized in that The computer-readable storage medium comprises instructions that, when the computer-readable storage medium is run on a server, cause the server to execute the method of any one of claims 10-17 and 40-46.
64. A computer program product comprising instructions, characterized in that When the computer program product is run on a server, the server is caused to execute the method of any one of claims 10-17 and 40-46.
65. A method for determining a posture, characterized in that The method comprises: Obtaining a pre-acquired reference image; For each reference image, determining a text region image contained in the reference image by performing text detection processing on the reference image; determining a text field contained in the text region image; Determining a 2D-3D correspondence relationship of each text region image based on the 2D points of the text region image contained in each reference image and a pre-acquired 2D-3D correspondence relationship of each reference image; The text fields and the reference images are stored in correspondence, and the 2D-3D correspondence relationship of the text area images is stored.
Citation Information
Patent Citations
Visual positioning method and device
CN109919157A
Location of image capture device and object features in a captured image
US20130243250A1