Data processing method and apparatus
By filtering and generating images similar to or related to the target scene in the navigation scenario, and using the NERF model for view localization, the problem of high view localization complexity in existing technologies is solved, achieving higher accuracy and wider view localization.
Patent Information
- Application Number
- CN202310768998.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-27
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2043-06-27
AI Technical Summary
In existing navigation scenarios, the perspective positioning scheme is highly complex and requires the use of satellites or other methods for positioning.
By acquiring the first image in the target scene, multiple images with high similarity or good viewpoint relationship are selected. More images are generated using the Neural Radiation Field (NERF) model or a 3D model, and viewpoint localization is performed based on the viewpoint data.
It reduces the complexity of viewpoint positioning and improves the accuracy and breadth of positioning.
Smart Images

Figure CN116823932B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and in particular to a data processing method and device. BACKGROUND
[0002] At present, in a navigation scene, a view angle needs to be positioned by means of satellites and the like.
[0003] Therefore, in the current positioning scheme, there is a defect of high complexity. SUMMARY
[0004] Therefore, the present application provides a data processing method and device, as follows:
[0005] A data processing method comprises the following steps:
[0006] obtaining a first image collected from a target scene;
[0007] obtaining a plurality of second images, the second images being images corresponding to the target scene;
[0008] determining a second target image from the plurality of second images according to the first image;
[0009] obtaining a plurality of third images, the third images being images corresponding to the target scene, each of the third images satisfying a first view angle relationship with the second target image;
[0010] determining a third target image from the plurality of third images according to the first image;
[0011] determining view angle data corresponding to the first image according to view angle data of the third target image.
[0012] Preferably, the method comprises the following steps:
[0013] obtaining a plurality of initial images, the initial images being images corresponding to the target scene, the plurality of initial images comprising at least one of the following: a plurality of randomly determined images, or a plurality of images selected by preset information;
[0014] obtaining the plurality of second images according to the plurality of initial images and the first image.
[0015] Preferably, the method comprises the following steps:
[0016] determining similarity information between the first image and each of the second images, and obtaining the second target image according to the similarity information; or
[0017] Identifying object images in the first image and each of the second images, and obtaining the second target image according to matching between the object images in the first image and each of the second images.
[0018] The method, preferably, the obtaining of the plurality of third images comprises:
[0019] Determining a first view range according to the view data of the second target image;
[0020] Determining a plurality of first view data in the first view range;
[0021] Obtaining the plurality of third images according to the plurality of first view data.
[0022] The method, preferably, the plurality of second images are obtained according to a second view range, and the determining of the first view range according to the view data of the second target image comprises:
[0023] Determining the first view range from the second view range according to the view data of the second target image, and the first view range is contained in the second view range.
[0024] The method, preferably, the plurality of second images are obtained according to a second view range, and the determining of the first view range according to the view data of the second target image comprises:
[0025] Determining that the view data of the second target image and a boundary of the second view range satisfy a first distance condition;
[0026] Determining the first view range according to the view data of the second target image, and the first view range partially overlaps with the second view range.
[0027] The method, preferably, the obtaining of the plurality of third images comprises:
[0028] Determining a plurality of first view data according to the view data of the second target image;
[0029] Inputting the plurality of first view data into a first model to obtain the plurality of third images, and the first model is trained to output a view image corresponding to view data in a target scene when receiving the view data in the target scene.
[0030] The method, preferably, the second image and the third image are obtained in the same way, and the number of the third images is greater than the number of the second images.
[0031] The method, preferably, the method further comprises at least one of the following:
[0032] In a case where the perspective data of the second target image and the perspective data of the first image satisfy a first angle relationship, the perspective data of the third image and the perspective data of the second target image satisfy a second angle relationship; or,
[0033] In a case where the perspective data of the second target image and the perspective data of the first image satisfy a first position relationship, the perspective data of the third image and the perspective data of the second target image satisfy a second position relationship.
[0034] A data processing apparatus comprises:
[0035] A first obtaining unit is configured to obtain a first image collected from a target scene;
[0036] A second obtaining unit is configured to obtain a plurality of second images, the second images being images corresponding to the target scene;
[0037] A target obtaining unit is configured to determine a second target image from the plurality of second images according to the first image;
[0038] A third obtaining unit is configured to obtain a plurality of third images, the third images being images corresponding to the target scene, and each of the third images satisfying a first perspective relationship with the second target image;
[0039] The target obtaining unit is further configured to determine a third target image from the plurality of third images according to the first image;
[0040] A perspective determining unit is configured to determine perspective data corresponding to the first image according to perspective data of the third target image.
[0041] As can be seen from the above technical solutions, in the data processing method and apparatus disclosed in the present application, after a first image in a target scene is collected, a plurality of images corresponding to the target scene are obtained, a corresponding target image is selected therefrom, and then a plurality of images corresponding to the target scene are obtained again based on a first perspective relationship, a corresponding target image is selected again, and then perspective data of the selected target image is determined as perspective data of the first image. It can be seen that, unlike the scheme of realizing perspective positioning by means of satellites in the prior art, the target image is obtained by selecting a plurality of images corresponding to the target scene in the present application, and then perspective data of the collected image in the target scene is positioned based on perspective data of the target image, so as to achieve the purpose of reducing positioning complexity. BRIEF DESCRIPTION OF DRAWINGS
[0042] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.
[0043] Figure 1 A flowchart of a data processing method provided in Embodiment One of the present application;
[0044] Figure 2 and Figure 3 are respectively example diagrams of the first view angle range and the second view angle range in the embodiments of the present application;
[0045] Figure 4 Another flowchart of a data processing method provided in Embodiment One of the present application;
[0046] Figure 5 A structural schematic diagram of a data processing device provided in Embodiment Two of the present application;
[0047] Figure 6 A structural schematic diagram of an electronic device provided in Embodiment Three of the present application;
[0048] Figure 7 A flowchart of NERF mapping in the present application;
[0049] Figure 8 A flowchart of positioning in a walking street scene L in the present application. DETAILED DESCRIPTION
[0050] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present application.
[0051] Reference Figure 1 As shown in the figure, an implementation flowchart of a data processing method provided in Embodiment One of the present application, which can be applied in an electronic device capable of image processing, such as a computer or a server, etc. The technical solution in the present embodiment is mainly used to realize view angle positioning of an image and reduce the complexity of view angle positioning.
[0052] Specifically, the method in the present embodiment can include the following steps:
[0053] Step 101: Obtain a first image collected from a target scene.
[0054] The target scene can be an indoor scene or an outdoor scene, such as an office, a shopping mall, or a pedestrian street.
[0055] Specifically, in this embodiment, the first image can be collected from the target scene by an image collection device. Alternatively, in this embodiment, the first image can be read from an image stored in a database, and the database stores at least one image collected from the target scene.
[0056] For example, in this embodiment, the first image is collected in a shopping mall by an image collection device such as a camera.
[0057] Step 102: Obtain a plurality of second images, the second images being images corresponding to the target scene.
[0058] It should be noted that the second image is an image corresponding to the target scene but not obtained by image collection for the target scene.
[0059] In one implementation, the target scene is pre-constructed with a neural radiance field (NERF) model, and the second image in this embodiment can be an image obtained based on the NERF model. The NERF model is obtained by collecting a sequence of images of the target scene and training.
[0060] In another implementation, the target scene is pre-constructed with a three-dimensional model, and the second image in this embodiment is an image obtained based on the three-dimensional model.
[0061] In another implementation, the target scene is pre-constructed with a gallery, and the gallery stores a plurality of images corresponding to the target scene. The second image in this embodiment is an image read from the gallery.
[0062] Step 103: Determine a second target image from the plurality of second images according to the first image.
[0063] The second target image is an image in the second images that satisfies a target similarity condition with the first image.
[0064] In one implementation, step 103 can first determine similarity information of the first image with each second image, and then obtain the second target image according to the similarity information.
[0065] For example, the first image and each second image are compared in similarity according to pixel values on corresponding pixel points, thereby obtaining similarity information of the first image with each second image. Then, the second image with similarity information greater than or equal to a similarity threshold is determined as the second target image.
[0066] For example, the image features extracted from the first image are compared with the image features extracted from each of the second images respectively to obtain similarity information of the first image and each of the second images, and then the second image with similarity information greater than or equal to a similarity threshold is determined as the second target image.
[0067] For example, similarity information of the first image and each of the second images is obtained by a similarity recognition model, where the similarity recognition model is trained with two input images as input and similarity label values as output, the similarity label values being the similarity information between the two input images, and then the second image with similarity information greater than or equal to a similarity threshold is determined as the second target image.
[0068] In another implementation, step 103 can first identify object images in the first image and each of the second images, and then obtain the second target image according to matching conditions between the object images in the first image and each of the second images.
[0069] The matching conditions represent similarity information between the object images.
[0070] For example, in step 103, object images in the first image and each of the second images are identified by an image recognition algorithm respectively, and then the object images in the first image are matched with the object images in each of the second images respectively to obtain matching conditions of the first image and each of the second images with respect to the object images, the matching conditions representing similarity information of the first image and the second image on the object images, and finally, the second image with the matching conditions satisfying a target similarity condition, such as similarity information greater than or equal to a similarity threshold, is determined as the second target image.
[0071] Step 104: Obtain a plurality of third images, the third images being images corresponding to the target scene, and each third image satisfying a first view angle condition with the second target image.
[0072] The first view angle condition means that the view angle data of the third image and the view angle data of the second target image satisfy a view angle proximity condition, which can be that the difference between the view angle data is less than or equal to a corresponding difference threshold, or the area of the view angle range in which the view angle data is located is less than or equal to a corresponding area threshold, and the like.
[0073] In an implementation, in the step 104 of obtaining the plurality of third images, the first view range can be determined according to the view angle data of the second target image first; then, in the first view range, a plurality of first view angle data is determined, for example, a plurality of view angle data is randomly selected in the first view range as the first view angle data, and the first view angle data is uniformly distributed in the first view range; finally, the plurality of third images is obtained according to the plurality of first view angle data, and the third image is one-to-one mapped with the first view angle data, that is, each third image corresponds to one of the first view angle data.
[0074] Optionally, the plurality of second images in the step 102 is obtained according to a second view range, based on which, in the step 104 of determining the first view range according to the view angle data of the second target image, the first view range can be determined from the second view range according to the view angle data of the second target image, and at this time, the first view range is contained in the second view range. For example, the first view range is a view range with the view angle data of the second target image as the center and a radius of a, as shown in FIG. 2. Figure 2
[0075] Optionally, the plurality of second images in the step 102 is obtained according to a second view range, based on which, in the step 104 of determining the first view range according to the view angle data of the second target image, the first view range can be determined from the second view range according to the view angle data of the second target image, and at this time, the first view range is contained in the second view range. For example, the first view range is a view range with the view angle data of the second target image as the center and a radius of a, as shown in FIG. 2. Figure 3
[0076] It should be noted that if the view angle data of the second target image does not satisfy the first distance condition with the boundary of the second view range, that is, the minimum difference between the view angle data of the second target image and the view angle data on the boundary of the second view range is greater than the boundary threshold, that is, the view range of the second target image is far away from the boundary of the second view range, then the first view range can be determined from the second view range according to the view angle data of the second target image, and at this time, the first view range is contained in the second view range. For example, the first view range is a view range with the view angle data of the second target image as the center and a radius of a, as shown in FIG. 2. Figure 2
[0077] In another implementation, in the step 104 of obtaining the plurality of third images, the plurality of first view data can be determined according to the view data of the second target image, and the first view data can be obtained as described above. Then, the plurality of first view data is input into the first model to obtain the plurality of third images. The first model is trained to output the view image corresponding to the view data in the target scene when receiving the view data in the target scene. For example, the first model can be a three-dimensional model trained for the target scene, which can output the view image corresponding to the first view data, i.e., the third image, after receiving the first view data. Alternatively, the first model can be a NERF model constructed for the target scene, which can output the view image corresponding to the first view data, i.e., the third image, after receiving the first view data.
[0078] In another implementation, in the case that the view data of the second target image and the view data of the first image satisfy the first angle relationship, the view data of the third image obtained in the step 104 and the view data of the second target image satisfy the second angle relationship.
[0079] In another implementation, whether the view data of the second target image and the view data of the first image satisfy the first angle relationship can be determined by the following method:
[0080] In another implementation, the image content analysis is performed on the second target image and the first image. In the case that the image content of the second target image and the image content of the first image satisfy the first similarity condition, the view data of the second target image and the view data of the first image satisfy the first angle relationship. In the case that the image content of the second target image and the image content of the first image do not satisfy the first similarity condition, the view data of the second target image and the view data of the first image do not satisfy the first angle relationship.
[0081] The first similarity condition can be that the similarity between the plurality of image contents recognized by the second target image and the plurality of image contents recognized by the first image is greater than or equal to the content similarity threshold, the poses of these image contents in the second target image are the same as the poses of these image contents in the first image, and the image regions of these image contents in the second target image are different from the image regions of these image contents in the first image. That is, the same image contents contained in the second target image and the first image exceed the first quantity threshold, and the poses of these same image contents in the second target image and the first image are the same, but the sizes are different. For example, the second target image and the first image both contain a building, and the orientations of the building in the second target image and the first image are the same, but the sizes are different.
[0082] Based on this, the second angle relationship matches the first angle relationship. That is, in the embodiment, when the third image is obtained, the view angle data can be kept consistent with the view angle data of the second target image, and multiple positions are selected and the corresponding view angle images at the positions are selected as the third image. For example, the multiple third images at the same view angle correspond to different shooting positions.
[0083] In another implementation, in a case where the view angle data of the second target image and the view angle data of the first image satisfy the first position relationship, the view angle data of the third image and the view angle data of the second target image satisfy the second position relationship.
[0084] The first position relationship between the view angle data of the second target image and the view angle data of the first image can be determined in the following manner:
[0085] The second target image and the first image are subjected to image content analysis, in a case where the image content of the second target image and the image content of the first image satisfy a second similarity condition, the view angle data of the second target image and the view angle data of the first image satisfy the first position relationship, and in a case where the image content of the second target image and the image content of the first image do not satisfy the second similarity condition, the view angle data of the second target image and the view angle data of the first image do not satisfy the first position relationship.
[0086] The second similarity condition can be that the similarity between the multiple image contents recognized by the second target image and the multiple image contents recognized by the first image is greater than or equal to a content similarity threshold value, the image contents are the same in the image area in the second target image and the image area in the first image, and the poses of the image contents in the second target image and the first image are different. That is, the same image contents contained in the second target image and the first image exceed the second quantity threshold value, and the same image contents are different in size but different in pose in the second target image and the first image, for example, a building is contained in the second target image and the first image, the size of the building in the second target image and the first image is the same, but the orientation is different.
[0087] Based on this, the second position relationship matches the first position relationship. That is, in the embodiment, when the third image is obtained, the position can be kept consistent with the position corresponding to the second target image, multiple view angle data are selected, and the corresponding view angle images at the view angle data are selected as the third image. For example, the multiple third images at the same position correspond to different shooting angles.
[0088] Step 105: determining a third target image from the multiple third images according to the first image.
[0089] The third target image is an image in the third images that meets a target similarity condition with the first image.
[0090] In an implementation, the step 105 can first determine similarity information of the first image with each of the third images, and then obtain the third target image according to the similarity information.
[0091] For example, the first image is compared with each of the third images in similarity according to pixel values on corresponding pixel points, so as to obtain similarity information of the first image with each of the third images, and then the third image with similarity information greater than or equal to a similarity threshold is determined as the third target image.
[0092] For another example, image features extracted from the first image are compared with image features extracted from each of the third images in similarity, so as to obtain similarity information of the first image with each of the third images, and then the third image with similarity information greater than or equal to a similarity threshold is determined as the third target image.
[0093] For another example, similarity information of the first image with each of the third images is obtained through a similarity recognition model, and then the third image with similarity information greater than or equal to a similarity threshold is determined as the third target image.
[0094] In another implementation, the step 105 can first recognize object images in the first image and each of the third images, and then obtain the third target image according to matching conditions between the object images in the first image and each of the third images.
[0095] The matching condition represents similarity information between the object images.
[0096] For example, object images in the first image and object images in each of the third images are recognized through an image recognition algorithm in the step 105, then the object images in the first image are matched with the object images in each of the third images, so as to obtain matching conditions of the first image with each of the third images about the object images, the matching conditions represent similarity information of the first image with the third images on the object images, and finally, the third image with the matching condition meeting a target similarity condition such as similarity information greater than or equal to a similarity threshold is determined as the third target image.
[0097] Step 106: determining view angle data corresponding to the first image according to view angle data of the third target image.
[0098] In an implementation, the view angle data of the third target image can be used as the view angle data of the first image.
[0099] In another implementation, the perspective data of the third target image can be adjusted according to matching of the third target image and the first image on the object image, and the adjusted perspective data is taken as the perspective data corresponding to the first image.
[0100] For example, the third target image and the first image both contain a certain building, the size of the building in the third target image and in the first image is the same, and only a slight difference in the pose exists, based on which, the perspective data of the third target image is slightly adjusted according to the difference in the pose, and thus the perspective data of the first image is obtained.
[0101] From the above technical solution, it can be seen that in the data processing method provided by the embodiment of the present application, after the first image in the target scene is collected, a plurality of images corresponding to the target scene are obtained, the corresponding target image is selected therefrom, and then the images corresponding to the target scene are re-obtained based on the first perspective relationship, the corresponding target image is selected again, and the perspective data of the selected target image is determined as the perspective data of the first image. It can be seen that, different from the scheme of realizing perspective positioning by means of satellites and the like in the prior art, in the embodiment, the target image is obtained through selection of a plurality of images corresponding to the target scene, and then the perspective data of the image collected in the target scene is positioned based on the perspective data of the target image, so as to achieve the purpose of reducing the positioning complexity.
[0102] It should be noted that after the third target image is determined in step 105, the third target image can be taken as the second target image, and step 104 is returned to be executed until the iteration termination condition is met, such as the iteration number exceeding a threshold value or the third target image meeting a target condition with the first image, such as shown in the following formula. Figure 4
[0103] The target condition here can be that the similarity information between the third target image and the first image is greater than or equal to a preset target threshold value, such as the similarity between the third target image and the first image being greater than 99.9%. Thus, through multiple iterations for multiple selections, the perspective data obtained in the last step 106 is more accurate.
[0104] In an implementation, when the plurality of second images are obtained in step 102, the following method can be used:
[0105] First, a plurality of initial images are obtained. The initial image is an image corresponding to the target scene. The initial image here can include at least one of the following:
[0106] A plurality of randomly determined images, for example, a plurality of images are randomly determined in the NERF model as initial images;
[0107] The plurality of images selected by the preset information, for example, in the NERF model, a plurality of images are selected as initial images according to the preset view angle data.
[0108] Then, a plurality of second images are obtained according to the plurality of initial images and the first image.
[0109] In an implementation manner, a plurality of images satisfying the initial condition with the first image can be directly screened from the plurality of initial images as the second images. The initial condition here can be that the similarity information is greater than or equal to the initial threshold. That is, a plurality of images whose similarity information with the first image is greater than or equal to the initial threshold are screened from the plurality of initial images as the second images. Then, in step 103, a second target image whose similarity information with the first image is greater than or equal to the similarity threshold is screened from the second images.
[0110] In another implementation manner, an image satisfying the target similarity condition with the first image can be first screened from the plurality of initial images as an intermediate image, and then a plurality of images satisfying the first view angle relationship with the intermediate image, i.e., the second images, are obtained according to the intermediate image.
[0111] The specific implementation manner of obtaining the intermediate image can refer to the manner of obtaining the second target image or the third target image in the foregoing. The manner of obtaining a plurality of images satisfying the first view angle relationship with the intermediate image can refer to the manner of obtaining the third image in the foregoing.
[0112] For example, the target scene is pre-built with a NERF model. In this embodiment, a plurality of initial images can be generated by the NERF model according to randomly selected view data, then the first image is used to filter out intermediate images from the initial images that meet the target similarity condition with the first image, then a plurality of view data are randomly selected from a view range centered on the view data of the intermediate image and with a radius of a, and then corresponding images are generated as second images according to the view data by the NERF model; then, second target images that meet the target similarity condition with the first image are filtered out from the second images; then, a plurality of view data are randomly selected from a view range centered on the view data of the second target image and with a radius of a, and then corresponding images are generated as third images according to the view data by the NERF model, and then third target images that meet the target similarity condition with the first image are filtered out from the third images; then, the third target image is taken as the second target image, and the process of randomly selecting a plurality of view data from a view range centered on the view data of the second target image and with a radius of a, and then generating corresponding images as third images according to the view data by the NERF model, and then filtering out third target images that meet the target similarity condition with the first image from the third images is repeated; and finally, the view data of the first image is determined according to the view data of the third target image.
[0113] It should be noted that the second images and the third images are obtained in the same way, and the number of the third images is greater than the number of the second images. Therefore, after the second target image is filtered out from the second images, more third images can be used for image filtering with higher view accuracy, so as to obtain the third target image, thereby making the view data corresponding to the first image more accurate.
[0114] Reference Figure 5 A structure schematic diagram of a data processing device provided in Embodiment Two of the present application is shown in the figure. The device can be configured in an electronic device capable of image processing, such as a computer or a server, etc. The technical solution in the present embodiment is mainly used for realizing view positioning of an image and reducing the complexity of view positioning.
[0115] Specifically, the device in the present embodiment can include the following units:
[0116] The first obtaining unit 501 is configured to obtain a first image collected from a target scene;
[0117] The second obtaining unit 502 is configured to obtain a plurality of second images, the second images being images corresponding to the target scene;
[0118] The target obtaining unit 503 is configured to determine a second target image from the plurality of second images according to the first image.
[0119] The third obtaining unit 504 is configured to obtain a plurality of third images, the third images being images corresponding to the target scene, and each of the third images satisfying a first view angle relationship with the second target image.
[0120] The target obtaining unit 503 is further configured to determine a third target image from the plurality of third images according to the first image.
[0121] The view angle determining unit 505 is configured to determine view angle data corresponding to the first image according to view angle data of the third target image.
[0122] As can be seen from the above technical solutions, in the data processing apparatus provided by Embodiment Two, after a first image in a target scene is collected, a plurality of images corresponding to the target scene are obtained, a corresponding target image is selected from the plurality of images, and then a plurality of images corresponding to the target scene are obtained again based on a first view angle relationship, a corresponding target image is selected again, and then view angle data of the selected target image is determined as view angle data of the first image. It can be seen that, unlike the scheme of realizing view angle positioning by means of satellites in the prior art, in this embodiment, a target image is obtained by selecting a plurality of images corresponding to the target scene, and then view angle data of an image collected in the target scene is positioned based on view angle data of the target image, so as to achieve the purpose of reducing positioning complexity.
[0123] In an implementation manner, the second obtaining unit 502 is specifically configured to: obtain a plurality of initial images, the initial images being images corresponding to the target scene, and the plurality of initial images including at least one of the following: a plurality of randomly determined images, or a plurality of images selected by preset information; and obtain the plurality of second images according to the plurality of initial images and the first image.
[0124] In an implementation manner, when the target obtaining unit 503 determines a second target image from the plurality of second images according to the first image, the target obtaining unit 503 includes at least one of the following:
[0125] determining similarity information of the first image and each of the second images, and obtaining the second target image according to the similarity information; or
[0126] identifying object images in the first image and each of the second images, and obtaining the second target image according to matching conditions between the object images in the first image and each of the second images, respectively.
[0127] In an implementation manner, the third obtaining unit 504 is specifically configured to: determine a first view angle range according to the view angle data of the second target image; determine a plurality of first view angle data in the first view angle range; and obtain the plurality of third images according to the plurality of first view angle data.
[0128] In an implementation manner, the plurality of second images are obtained according to a second view angle range, and the third obtaining unit 504 is specifically configured to: determine the first view angle range from the second view angle range according to the view angle data of the second target image, and the first view angle range is contained in the second view angle range when determining the first view angle range according to the view angle data of the second target image.
[0129] In an implementation manner, the plurality of second images are obtained according to a second view angle range, and the third obtaining unit 504 is specifically configured to: determine that the view angle data of the second target image satisfies a first distance condition with a boundary of the second view angle range; and determine the first view angle range according to the view angle data of the second target image, and the first view angle range partially overlaps with the second view angle range.
[0130] In an implementation manner, the third obtaining unit 504 is specifically configured to: determine a plurality of first view angle data according to the view angle data of the second target image; and input the plurality of first view angle data into a first model to obtain the plurality of third images, and the first model is trained to output a view angle image corresponding to the view angle data in the target scene when receiving the view angle data in the target scene.
[0131] In an implementation manner, the second image and the third image are obtained in the same manner, and the number of the third images is greater than the number of the second images.
[0132] In an implementation manner, the view angle data of the third image satisfies a second angle relationship with the view angle data of the second target image when the view angle data of the second target image satisfies a first angle relationship with the view angle data of the first image; or,
[0133] In an implementation manner, the view angle data of the third image satisfies a second position relationship with the view angle data of the second target image when the view angle data of the second target image satisfies a first position relationship with the view angle data of the first image.
[0134] It should be noted that the specific implementation of each unit in the present embodiment can refer to the corresponding content in the foregoing, which will not be described in detail here.
[0135] Reference Figure 6A structure diagram of an electronic device is provided for Embodiment Three of the present application. The electronic device can be a device capable of data processing, such as a computer or a server, etc. The electronic device can include the following structure:
[0136] The memory 601 is configured to store computer programs and data generated during the running of the computer programs.
[0137] The processor 602 is configured to execute the computer programs to achieve the following: obtaining a first image collected from a target scene; obtaining a plurality of second images corresponding to the target scene; determining a second target image from the plurality of second images according to the first image; obtaining a plurality of third images corresponding to the target scene, each of the third images satisfying a first view angle relationship with the second target image; determining a third target image from the plurality of third images according to the first image; and determining view angle data corresponding to the first image according to view angle data of the third target image.
[0138] As can be seen from the above technical solution, in the electronic device of Embodiment Three of the present application, after a first image in a target scene is collected, a plurality of images corresponding to the target scene are obtained, and a corresponding target image is selected therefrom, and then a plurality of images corresponding to the target scene are obtained again based on a first view angle relationship, and a corresponding target image is selected again, and then view angle data of the selected target image is determined as view angle data of the first image. It can be seen that, unlike the scheme of realizing view angle positioning by means of satellites in the prior art, in the present embodiment, the target image is obtained by selecting a plurality of images corresponding to the target scene, and then the view angle data of the image collected in the target scene is positioned based on the view angle data of the target image, thereby achieving the purpose of reducing the positioning complexity.
[0139] Taking the scene of a pedestrian street as an example, the technical solution in the present embodiment is described as follows:
[0140] First, a sequence of images is collected in advance, a NERF file of the current scene is generated by NERF training, and pictures and poses of sparse view angles are generated in the NERF for positioning.
[0141] Second, when view angle positioning is needed, a similar comparison is made between the current picture and the picture under the sparse view angle, a new sparse picture is regenerated by NERF within the nearby view angle range of the most similar sparse picture, until the iteration is terminated (such as the number of iterations reaches a threshold or the similarity between the selected picture and the current picture is high), and finally the pose corresponding to the most similar image is output.
[0142] Reference Figure 7 The flowchart for NERF mapping: first, a sequence of images is collected in the current scene L, such as set A (a1, a2, a3, …, an), and then the NERF file is generated by NERF training.n Then, set A is trained in NERF to obtain the NERF file (NERF.file); finally, some images, such as set B (b1, b2, b3, ..., b...), can be uniformly generated through NERF.file. n Each image corresponds to a pose (i.e., viewpoint), such as a set P(p1, p2, p3, ..., p...). n ).
[0143] refer to Figure 8 The flowchart for the localization process in this application is as follows: An image H is acquired in the current scene L; then, a similarity judgment is made between H and the image in B, and the most similar image b is obtained in B. i p in P i Next, determine whether the number of iterations has reached the threshold or the similarity value has reached the similarity threshold. If not, proceed to p. i Within a radius of α, generate new discrete images using NERF.file, replacing the data in sets B and P; then return to find the most similar image in B. i p in P i The process continues until the number of iterations reaches a threshold or the similarity value reaches a similarity threshold, at which point p is output. i As the pose of H in NERF.file.
[0144] As can be seen, this embodiment uses NERF mapping and an iterative method based on image comparison to achieve localization. Therefore, the technical solution of this embodiment can achieve higher accuracy localization and a wider localization perspective.
[0145] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.
[0146] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0147] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and
[0148] The above description of disclosed embodiments is intended to be illustrative and not restrictive. Many embodiments of the application will be apparent to those of skill in the art upon reviewing the above description. The scope of the application should, therefore, be determined not with reference to the above description, but instead should be given with reference to the appended claims, along with their full scope of equivalents.
Claims
1. A data processing method comprising: obtaining a first image collected from a target scene; obtaining a plurality of second images, the second images being images corresponding to the target scene; determining a second target image from the plurality of second images according to the first image, the second target image being an image in the second images that satisfies a target similarity condition with the first image; obtaining a plurality of third images, the third images being images corresponding to the target scene, each of the third images satisfying a first view angle relationship with the second target image; determining a third target image from the plurality of third images according to the first image, the third target image being an image in the third images that satisfies a target similarity condition with the first image; determining view angle data corresponding to the first image according to view angle data of the third target image; wherein the obtaining of the plurality of third images comprises: determining a first view angle range according to the view angle data of the second target image; determining a plurality of first view angle data in the first view angle range; and obtaining the plurality of third images according to the plurality of first view angle data.
2. A data processing method comprising: obtaining a first image collected from a target scene; obtaining a plurality of second images, the second images being images corresponding to the target scene; determining a second target image from the plurality of second images according to the first image, the second target image being an image in the second images that satisfies a target similarity condition with the first image; obtaining a plurality of third images, the third images being images corresponding to the target scene, each of the third images satisfying a first view angle relationship with the second target image; determining a third target image from the plurality of third images according to the first image, the third target image being an image in the third images that satisfies a target similarity condition with the first image; determining view angle data corresponding to the first image according to view angle data of the third target image; wherein the obtaining of the plurality of third images comprises: determining a plurality of first view angle data according to the view angle data of the second target image; and inputting the plurality of first view angle data into a first model to obtain the plurality of third images, the first model being trained to output a view angle image corresponding to view angle data in the target scene when receiving the view angle data in the target scene.
3. The method of claim 1 or 2, wherein the obtaining of the plurality of second images comprises: obtaining a plurality of initial images, the initial images being images corresponding to the target scene, the plurality of initial images comprising at least one of: a plurality of randomly determined images, or a plurality of images selected by preset information; and obtaining the plurality of second images according to the plurality of initial images and the first image.
4. The method of claim 1 or 2, wherein the determining of the second target image from the plurality of second images according to the first image comprises at least one of: determining similarity information of the first image and each of the second images, and obtaining the second target image according to the similarity information; or identifying object images in the first image and each of the second images, and obtaining the second target image according to matching between the object images in the first image and each of the second images.
5. The method of claim 1 or 2, wherein the second images are obtained according to a second view range, and the first view range is determined according to view data of the second target image, including: determining the first view range from the second view range according to the view data of the second target image, the first view range being contained in the second view range.
6. The method of claim 1 or 2, wherein the second images are obtained according to a second view range, and the first view range is determined according to view data of the second target image, including: determining that the view data of the second target image and a boundary of the second view range satisfy a first distance condition; and determining the first view range according to the view data of the second target image, the first view range partially overlapping with the second view range.
7. The method of claim 1 or 2, wherein the second images and the third images are obtained in the same way, and the number of the third images is greater than the number of the second images.
8. The method of claim 1 or 2, further comprising at least one of: in a case where the view data of the second target image and the view data of the first image satisfy a first angle relationship, the view data of the third image and the view data of the second target image satisfy a second angle relationship; or, in a case where the view data of the second target image and the view data of the first image satisfy a first position relationship, the view data of the third image and the view data of the second target image satisfy a second position relationship.
9. A data processing apparatus, comprising: a first obtaining unit configured to obtain a first image collected from a target scene; a second obtaining unit configured to obtain a plurality of second images, the second images being images corresponding to the target scene; a target obtaining unit configured to determine a second target image from the plurality of second images according to the first image, the second target image being an image in the second images satisfying a target similarity condition with the first image; a third obtaining unit configured to obtain a plurality of third images, the third images being images corresponding to the target scene, each of the third images satisfying a first view relationship with the second target image; the target obtaining unit is further configured to determine a third target image from the plurality of third images according to the first image, the third target image being an image in the third images satisfying the target similarity condition with the first image; a view determining unit configured to determine view data corresponding to the first image according to view data of the third target image; and wherein the obtaining of the plurality of third images comprises: determining a first view range according to the view data of the second target image; determining a plurality of first view data in the first view range; and obtaining the plurality of third images according to the plurality of first view data.
10. A data processing apparatus, comprising: The first obtaining unit is configured to obtain a first image collected from a target scene; The second obtaining unit is configured to obtain a plurality of second images, the second images being images corresponding to the target scene; The target obtaining unit is configured to determine a second target image from the plurality of second images according to the first image, the second target image being an image in the second images that satisfies a target similarity condition with the first image; The third obtaining unit is configured to obtain a plurality of third images, the third images being images corresponding to the target scene, and each of the third images satisfying a first view angle relationship with the second target image; The target obtaining unit is further configured to determine a third target image from the plurality of third images according to the first image, the third target image being an image in the third images that satisfies a target similarity condition with the first image; The view angle determining unit is configured to determine view angle data corresponding to the first image according to view angle data of the third target image; The plurality of third images are obtained by: determining a plurality of first view angle data according to the view angle data of the second target image; and inputting the plurality of first view angle data into a first model to obtain the plurality of third images, the first model being trained to, in a case where view angle data in the target scene is received, output a view angle image corresponding to the view angle data in the target scene.
Citation Information
Patent Citations
Method for visual localization and related apparatus
US20220148302A1