Image data processing method, device, storage medium and program product
By acquiring image and pose information, using a prediction model to determine the depth information of three-dimensional spatial points and distinguishing between scanned and unscanned points, the problem of insufficient image completeness in online painting purchases is solved, and the user interaction experience is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ALIBABA (CHINA) CO LTD
- Filing Date
- 2022-11-28
- Publication Date
- 2026-05-08
AI Technical Summary
Replacing traditional decorative paintings requires users to select them in person, which is time-consuming and laborious. Online painting purchases cannot guarantee the completeness of the photos and have low user interaction, requiring users to manually fill in any missing areas.
By acquiring image and pose information of the target spatial region, the depth information of three-dimensional spatial points is determined using a prediction model, and the scanned and unscanned points are displayed in different styles on the interactive interface, reducing the need for manual judgment by users.
It improves the terminal's interactive performance, allowing users to determine the depth information of three-dimensional spatial points in the image without inputting a depth map, automatically prompting for missed areas, and simplifying the process of selecting decorative paintings.
Smart Images

Figure CN115731293B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to an image data processing method, device, storage medium, and program product. Background Technology
[0002] In the decorative painting industry, annual consumption nationwide is approximately 500 billion yuan, with an average of 4-5 paintings hung per 100 square meters. As people's living standards improve, the demand for replacing decorative paintings is increasing. Traditionally, replacing decorative paintings requires customers to go to the decorative painting market to select them, which is time-consuming and laborious.
[0003] With the development of information technology, online art purchasing has become a new choice due to its convenience and speed. Online art purchasing platforms typically display images and descriptions of available decorative paintings for users to choose from. During the selection process, users can take photos of potential hanging areas, such as interior walls, and upload them to the platform. This process requires users to determine suitable hanging locations, remember the specific areas photographed, and manually fill in any missing areas. The entire process not only lacks completeness in terms of image quality but also suffers from low user interaction and is time-consuming and laborious. Summary of the Invention
[0004] The main objective of this application is to provide an image data processing method, device, storage medium, and program product that enables the acquisition of depth information of three-dimensional spatial points in an image without the need for an input depth map. The scanned points are displayed in a preset style to indicate to the user which areas were missed, eliminating the need for manual user judgment and improving terminal interaction performance.
[0005] In a first aspect, embodiments of this application provide an image data processing method, comprising: responding to a user's first operation on a target spatial region, acquiring first image information of the target spatial region and first pose information corresponding to the first image information, wherein the first image information includes at least: a group of pixels obtained by capturing images of some three-dimensional spatial points in the target spatial region; determining, based on the first pose information and the first image information, position estimation information of each three-dimensional spatial point corresponding to each pixel in the first image information in the target spatial region; determining, based on the first image information and the first pose information, distance information between each three-dimensional spatial point in the target spatial region and the surface of a target object; determining, based on the distance information and the position estimation information, depth information of each three-dimensional spatial point in the target spatial region; and displaying, based on the depth information, the pixels corresponding to each three-dimensional spatial point on an interactive interface in a first style.
[0006] In one embodiment, determining the position estimation information of the three-dimensional spatial points corresponding to each pixel in the first image information in the target spatial region based on the first pose information and the first image information includes: performing three-dimensional projection on each pixel in the first image information based on the first pose information, and using the obtained three-dimensional projection points as the position estimation information of each three-dimensional spatial point in the target spatial region.
[0007] In one embodiment, determining the distance information between each three-dimensional spatial point in the target space region and the surface of the target object based on the first image information and the first pose information includes: inputting the first image information and the first pose information into a preset prediction model, and outputting the distance information between the three-dimensional spatial point corresponding to each pixel in the first image information and the surface of the target object in the target space region.
[0008] In one embodiment, the step of training the prediction model includes: acquiring multi-channel features of a sample image and pose information of a preset frame image in the sample image; mapping the multi-channel features to a three-dimensional feature space based on the pose information of the preset frame image to obtain a three-dimensional feature map of the sample image; and training a preset neural network using the three-dimensional feature map to obtain the prediction model.
[0009] In one embodiment, determining the depth information of each three-dimensional spatial point in the target spatial region based on the distance information and the position estimation information includes: determining the vector difference between the vector of the position estimation information and the vector of the distance information, and using the obtained vector difference as the depth information of each three-dimensional spatial point in the target spatial region.
[0010] In one embodiment, the method further includes: in response to a second operation by a user on the target spatial region, acquiring second image information of the target spatial region and second pose information corresponding to the second image information; determining, based on the second image information and the second pose information, the depth information of the three-dimensional spatial points corresponding to each pixel to be retrieved in the second image information in the target spatial region; determining the scanning state of each pixel to be retrieved based on the depth information, and distinguishing between scanned and unscanned points in the second image information on the interactive interface.
[0011] In one embodiment, determining the scanning status of each pixel to be retrieved in the second image information based on the depth information, and distinguishing between scanned and unscanned points in the second image information on the interactive interface, includes: determining whether each pixel to be retrieved in the second image information is stored in the scanned point set based on the depth information; and displaying the first pixel to be retrieved in the second image information that is stored in the scanned point set on the interactive interface using a first style.
[0012] In one embodiment, the step of determining the scanning status of each pixel to be retrieved in the second image information based on the depth information and distinguishing between scanned and unscanned points in the second image information on the interactive interface further includes: displaying second pixels to be retrieved in the second image information that are not stored in the set of scanned points using a second style on the interactive interface, wherein the second style is different from the first style.
[0013] In one embodiment, determining whether each pixel to be retrieved in the second image information is stored in the scanned point set based on the depth information includes: for each pixel to be retrieved in the second image information, searching the scanned point set according to the corresponding depth information; if a point with the same depth information as the current pixel to be retrieved is found, then it is determined that the current pixel to be retrieved is stored in the scanned point set; otherwise, it is determined that the current pixel to be retrieved is not stored in the scanned point set.
[0014] In one embodiment, the distinguishing display of scanned points and unscanned points includes: displaying the scanned points and unscanned points at corresponding positions on the image information.
[0015] In one embodiment, the method further includes: processing each pixel in the image information in parallel.
[0016] Secondly, embodiments of this application provide an image data processing method, including: in response to a user's shooting operation on an indoor scene, acquiring first image information of the indoor scene and first pose information corresponding to the first image information; performing three-dimensional reconstruction of the indoor scene based on the first image information and the first pose information; processing each pixel in the first image information using any of the methods described above, and distinguishing between scanned and unscanned points in the first image information.
[0017] Thirdly, embodiments of this application provide an image data processing apparatus, comprising:
[0018] The acquisition module is used to respond to the user's first operation on the target spatial region, acquire the first image information of the target spatial region and the first pose information corresponding to the first image information, wherein the first image information includes at least: a group of pixels obtained by taking pictures of some three-dimensional spatial points in the target spatial region.
[0019] The first determining module is used to determine the position estimation information of the three-dimensional spatial points corresponding to each pixel in the first image information in the target spatial region based on the first pose information and the first image information.
[0020] The second determining module is used to determine the distance information between each three-dimensional spatial point and the surface of the target object in the target spatial region based on the first image information and the first pose information.
[0021] The third determining module is used to determine the depth information of each three-dimensional spatial point in the target spatial region based on the distance information and the position estimation information;
[0022] The display module is used to display the pixels corresponding to each three-dimensional spatial point on the interactive interface in a first style according to the depth information.
[0023] In one embodiment, the first determining module is used to perform three-dimensional projection on each pixel in the first image information according to the first pose information, and the obtained three-dimensional projection points are used as position estimation information of each three-dimensional spatial point in the target spatial region.
[0024] In one embodiment, the second determining module is used to input the first image information and the first pose information into a preset prediction model, and output the distance information between each three-dimensional spatial point corresponding to each pixel in the first image information and the surface of the target object in the target space region.
[0025] In one embodiment, the system further includes a training module for training the prediction model, comprising: acquiring multi-channel features of a sample image and pose information of a preset frame image in the sample image; mapping the multi-channel features to a three-dimensional feature space based on the pose information of the preset frame image to obtain a three-dimensional feature map of the sample image; and training a preset neural network using the three-dimensional feature map to obtain the prediction model.
[0026] In one embodiment, the third determining module is used to determine the vector difference between the vector of the position estimation information and the vector of the distance information, and the obtained vector difference is used as the depth information of each three-dimensional spatial point in the target spatial region.
[0027] In one embodiment, the system further includes: a retrieval module, configured to respond to a second operation by a user on a target spatial region, acquire second image information of the target spatial region and second pose information corresponding to the second image information; determine, based on the second image information and the second pose information, the depth information of the three-dimensional spatial points corresponding to each pixel to be retrieved in the second image information within the target spatial region; determine the scanning state of each pixel to be retrieved based on the depth information, and distinguish between scanned and unscanned points in the second image information on the interactive interface.
[0028] In one embodiment, the retrieval module is used to determine, based on the depth information, whether each pixel to be retrieved in the second image information is stored in the scanned point set; and to display the first pixel to be retrieved in the second image information that is stored in the scanned point set on the interactive interface using a first style.
[0029] In one embodiment, the retrieval module is further configured to display the second retrieval point in the second image information that is not stored in the scanned point set in a second style on the interactive interface, the second style being different from the first style.
[0030] In one embodiment, the retrieval module is further configured to, for each pixel to be retrieved in the second image information, retrieve it in the scanned point set according to the corresponding depth information. If a point with the same depth information as the current pixel to be retrieved is found, it is determined that the current pixel to be retrieved is stored in the scanned point set; otherwise, it is determined that the current pixel to be retrieved is not stored in the scanned point set.
[0031] In one embodiment, the distinguishing display of scanned points and unscanned points includes: displaying the scanned points and unscanned points at corresponding positions on the image information.
[0032] In one embodiment, the method further includes: processing each pixel in the image information in parallel.
[0033] Fourthly, embodiments of this application provide an electronic device, including:
[0034] At least one processor; and
[0035] A memory that is communicatively connected to the at least one processor;
[0036] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, cause the electronic device to perform the method described in any of the above aspects.
[0037] Fifthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method described in any of the above aspects.
[0038] Sixthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the above aspects.
[0039] The image data processing method, device, storage medium, and program product provided in this application embodiment obtain image information of a target spatial region and its corresponding pose information based on user operations. This information is used to determine the depth information of three-dimensional spatial points in the image information. The pixels corresponding to the three-dimensional spatial points with determined depth information are displayed on the interactive interface in a preset first style. In this way, the depth information of three-dimensional spatial points can be obtained without inputting a depth map. By displaying the pixels corresponding to the three-dimensional spatial points that have been scanned in the first image information in the first style, the user is prompted which points have been scanned, which helps the user to judge which areas of image information are incomplete. This eliminates the need for manual judgment by the user and improves the interactive performance of the terminal. Attached Figure Description
[0040] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are some embodiments of the invention, and that those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0041] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0042] Figure 2A This is a schematic diagram of an image data processing system provided in an embodiment of this application;
[0043] Figure 2B This is a schematic diagram illustrating an application scenario for image data processing provided in an embodiment of this application;
[0044] Figure 3 A flowchart illustrating an image data processing method provided in an embodiment of this application;
[0045] Figure 4 A flowchart illustrating an image data processing method provided in an embodiment of this application;
[0046] Figure 5 A flowchart illustrating an image data processing method provided in an embodiment of this application;
[0047] Figure 6 This is a schematic diagram of the structure of an image data processing device provided in an embodiment of this application.
[0048] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0049] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0050] In this article, the term "and / or" is used to describe the relationship between related objects. Specifically, it means that there can be three kinds of relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, or B exists alone.
[0051] To clearly describe the technical solutions of the embodiments of this application, the terms involved in this application are first defined as follows:
[0052] TSDF: Truncated Signed Distance Function. The TSDF of a point in three-dimensional space stores the distance from that point to the nearest surface of the object. Positive and negative values indicate whether the point is inside or outside the surface.
[0053] TSDF fusion: Truncation of directed range field fusion. By integrating different TSDF values, it draws contour surfaces corresponding to the zero values to fit the object surface.
[0054] FPN: Feature Pyramid Networks.
[0055] Scanning: refers to the process by which an image acquisition device acquires images of a target spatial area, such as the process of a camera taking a picture. The points that are accurately acquired are the scanned points, and the points that are not captured are the unscanned points.
[0056] like Figure 1 As shown, this embodiment provides an electronic device 1, including: at least one processor 11 and a memory 12. Figure 1Taking a processor as an example, the processor 11 and the memory 12 are connected via a bus 10. The memory 12 stores instructions that can be executed by the processor 11. The instructions are executed by the processor 11 to enable the electronic device 1 to perform all or part of the process of the method in the following embodiments, so as to determine the position of the object point in the target spatial region without the need to introduce an additional depth map, and display the scanned points in a preset style to prompt the user which areas were missed. This eliminates the need for manual judgment by the user and improves the computing performance and interactive performance of the terminal.
[0057] In one embodiment, the electronic device 1 may be a mobile phone, tablet computer, laptop computer, desktop computer, or a large computing system composed of multiple computers.
[0058] Figure 2A This is a schematic diagram of an image data processing system 200 provided in an embodiment of this application. Figure 2A As shown, the system includes: a server 210 and a terminal 220, wherein:
[0059] Server 210 can be a data platform providing data resource services, such as an online decorative painting purchasing platform. In a real-world scenario, a decorative painting purchasing platform may have multiple servers 210. Figure 2A Taking a single server (210) as an example.
[0060] Terminal 220 can be a computer, mobile phone, tablet, or other device used by the user to log in to the decorative painting purchasing platform. There can also be multiple terminals 220. Figure 2A The following example uses two terminals, 220, for illustration.
[0061] Terminal 220 and server 210 can transmit information via the Internet, enabling terminal 220 to access data on server 210. Both terminal 220 and / or server 210 can be implemented by electronic device 1.
[0062] The image data processing method of this application embodiment can be applied to any field that requires image data processing.
[0063] For example, the image data processing method of this application embodiment can be applied to scenarios such as online painting purchases or the purchase of furniture and other decorative items. Taking the online painting purchase scenario as an example, during the process of a user selecting a painting, the user needs to take pictures of possible hanging areas such as indoor walls, upload the pictures to the platform, and the user needs to determine the location where the painting can be hung during the shooting process. During this process, the user needs to remember the specific locations that have been photographed and determine whether all the pictures have been taken. If not, the user needs to manually fill in the missing areas. The entire process not only cannot guarantee the completeness of the shooting, but also has low user interaction and is time-consuming and laborious. In the currently available embodiments, an index + search method can generally be used to realize the reminder of the missed areas, and the specific process is as follows:
[0064] Indexing Phase: Based on the concept of point cloud fusion, for the scanned area in an indoor scene, the 2D points captured by the video are sequentially mapped to the camera coordinate system according to the given 2D coordinates and corresponding pose and depth maps. Subsequently, camera extrinsic parameters are used to transform the data to the world coordinate system, and the precise coordinates of the point on the corresponding object surface are calculated using the depth map. A matrix is used to record its position and color information. To achieve constant-time lookup operations, a structured hash table is used to record the points that have appeared. Subsequent lookup operations are also implemented by accessing the corresponding index in the structured hash table.
[0065] Search Phase: For the point to be queried, the point's pose parameters are used to project the point onto its corresponding position in 3D space. A structured hash table is accessed to determine if the point has been recorded. If so, the point is marked as scanned (displayed as pixels at this location in visualization); otherwise, it is an unscanned point (represented as a raster in visualization).
[0066] The disadvantages of the above-described embodiments are:
[0067] 1. Modeling large indoor spaces typically requires 30,000 to 100,000 vertices. Mapping all points in a new input image and then individually identifying each vertex is very time-consuming. Although structured hash tables are used for speed optimization, the effect is still slow, especially with high image resolution, where the speed cannot fully meet real-time requirements. Therefore, in practice, this solution updates every few frames, resulting in a certain delay in the 3D reconstruction and field scan completeness indicators.
[0068] 2. In the above solution, when users use the product, a real-time depth map of the current scene is needed to accurately estimate the position of the object's surface coordinates in the world coordinate system. This adds extra data requirements. Considering the actual configuration of most mid-range and low-end mobile phones, this requirement is difficult to meet.
[0069] 3. If step 2 cannot be satisfied, depth map prediction will also be considered to achieve this function. However, depth map prediction introduces an additional accuracy bottleneck, places additional demands on mobile phone performance, and also places higher requirements on product deployment.
[0070] In addition, wall line detection combined with corner detection can be used to determine the location of scanned walls indoors, and the completeness of scene scanning can be assessed by estimating the number and approximate location of scanned walls. However, since this method only uses key points such as wall corners for simple determination without involving 3D coordinate information, the accuracy of this approach is more difficult to guarantee (for example, if a user misses a point, the entire process will fail and cannot be closed). It requires robust interaction logic to be successfully implemented. Furthermore, since some tasks require 3D reconstruction, this method cannot be executed in parallel with related tasks, resulting in low efficiency.
[0071] Therefore, the aforementioned scheme combining wall line detection and corner detection cannot simultaneously support the smooth execution of the 3D reconstruction subtask, resulting in low coupling. Simply identifying corners may lead to inaccurate differentiation in special cases (such as floor-to-ceiling windows, beds, and doors), resulting in misclassification as wall corners, yet still classifying the scan as complete when it is not actually finished. This leads to poor overall system stability and increased uncertainty in practical use. Furthermore, because it only checks if a corner has been scanned, ignoring the assessment of all points within the target space area at the current scan position, the prompting algorithm cannot align with the 3D reconstruction algorithm. Ultimately, this results in poor 3D reconstruction results even though the scan completion prompting algorithm indicates completion. For example, it might determine that the wall has been scanned, but the sofa in front of the wall has not. The root cause of this situation is the failure to utilize 3D information (depth map) for point-by-point matching, relying solely on 2D vision to determine the completion of specific key points.
[0072] To address the aforementioned issues, this application provides an image data processing scheme that enables 3D spatial shooting completion indication based on video frame pose. Without using additional information (such as depth information), it accurately indicates the user's completed shooting locations (represented by bright colors) and incomplete shooting locations (represented by dark colors) in real-time scenes using a regular mobile phone. Users can then supplement the incomplete shooting portions of the corresponding scene according to the given results to complete the entire scene shooting process.
[0073] like Figure 2B The diagram shown illustrates an application scenario for image data processing provided in this application embodiment. Taking a 3D reconstruction scenario as an example, the input required for this embodiment can be real-time captured video information and pose information corresponding to each frame. Subsequent processing can proceed in two directions.
[0074] Option 1: Based on the input video information and the pose information corresponding to each frame, a TSDF prediction network is used to directly predict the TSDF value using camera video frames and pose information. Here, the TSDF value represents the depth difference between a 3D point position and its nearest surface in a video frame captured by the camera. The directly predicted TSDF value can help predict the 3D reconstruction result of the target spatial region, thus enabling direct 3D reconstruction.
[0075] Option 2: Based on the input video information and the pose information corresponding to each frame, perform pose and position estimation, i.e., pose projection, to obtain the estimated point coordinates (the corresponding result is the estimated 3D point projection). For the same point, further subtract the predicted TSDF value from the estimated point coordinates to obtain the estimated depth information value for that point. Based on this, the estimated depth information value of the point can be directly used as the accurate 3D point projection of that point, stored in the scanned point set, indicating that the point has been scanned, and distinguishing between scanned and unscanned points for display.
[0076] The aforementioned image data processing solution can be deployed on server 210, on terminal 220, or partially on server 210 and partially on terminal 220. The appropriate solution can be chosen based on actual needs in a real-world scenario; this embodiment does not impose any limitations.
[0077] When the image data processing solution is deployed entirely or partially on the server 210, the call interface can be opened to the terminal 220 to provide algorithm support to the terminal 220.
[0078] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Where there is no conflict between the embodiments, the following embodiments and features can be combined with each other. Furthermore, the timing of the steps in the following method embodiments is merely an example and not a strict limitation.
[0079] Please refer to Figure 3 This is an embodiment of an image data processing method according to this application. The method can be performed by... Figure 1 The electronic device 1 shown is used to perform this action and can be applied to... Figure 2A and Figure 2B In the application scenario of the image data processing system shown, the goal is to determine the position of object points in the target spatial region without the need for additional depth maps, and to display the scanned points in a preset first style to indicate to the user which areas were missed. This eliminates the need for manual user judgment, improving the terminal's computing and interactive performance. This embodiment uses terminal 220 as the execution end as an example, and the method includes the following steps:
[0080] Step 301: In response to the user's first operation on the target spatial region, obtain the first image information of the target spatial region and the first pose information corresponding to the first image information.
[0081] In this step, the first operation is used to initiate the image information acquisition process, such as turning on the camera. The first operation can be a shooting operation of a target spatial area, for example, a user opening their phone's camera and shooting the target spatial area. The target spatial area can be a user-specified area, an area of an indoor scene, such as a bedroom or living room, or an area containing multiple rooms. The first image information includes at least: a group of pixels obtained by shooting a portion of the three-dimensional spatial points in the target spatial area; that is, the pixels in the first image information correspond to the three-dimensional spatial points in the target spatial area. When a user selects a painting through an online art purchase platform, they can turn on their phone's camera and shoot the bedroom, acquiring the first image information of the indoor scene. This image information can be continuous video information or one or more unordered image information. Based on the basic parameters of the phone's camera, the first pose information corresponding to the first image information can be obtained. This pose information can be the camera's extrinsic parameters, such as the rotation angle. When the first image information is video information, the first pose information includes the pose information corresponding to each frame of the video information; when the first image information is one or more unordered image information, the first pose information includes the pose information corresponding to each image information.
[0082] Step 302: Based on the first pose information and the first image information, determine the position estimation information of each three-dimensional spatial point corresponding to each pixel in the first image information in the target spatial region.
[0083] In this step, the 3D spatial point can be any point in the target spatial region. Based on the first pose information and the first image information, the position of each 3D spatial point corresponding to each pixel in the first image information can be estimated. Position estimation can be accomplished through coordinate transformation. Assuming the 3D spatial point is a point in the bedroom, and the first pose information includes the position coordinates of the 3D spatial point in the camera coordinate system, the position of the 3D spatial point in the bedroom can be estimated based on the positional relationship between the 3D spatial point and other objects, thus obtaining the position estimation information.
[0084] In one embodiment, step 302 may specifically include: performing three-dimensional projection on each pixel in the first image information based on the first pose information, and using the obtained three-dimensional projection points as position estimation information of each three-dimensional spatial point in the target spatial region.
[0085] In this step, the first image information is two-dimensional image information. Its corresponding pose information includes camera pose information, such as rotation angle and translation parameters. Based on its corresponding first pose information, the projection point of each pixel in the two-dimensional image information in three-dimensional space can be predicted. This projection point serves as the position estimation information for each corresponding three-dimensional spatial point in the target spatial region. For each pixel in the first image information, based on the idea of point cloud fusion, for the scanned area in the indoor scene, according to the given two-dimensional coordinates and corresponding pose information, the two-dimensional pixels captured by the video are sequentially projected into three-dimensional space to obtain the position estimation information. Since the user does not need to input a depth map, using pose information to predict the three-dimensional position of each three-dimensional spatial point in the target spatial region can reduce the amount of data computation.
[0086] Step 303: Based on the first image information and the first pose information, determine the distance information between each three-dimensional spatial point and the surface of the target object in the target spatial region.
[0087] In this step, the target object can be a reference object within the target space region. For example, in selecting a scene for a bedroom, the target space region is the bedroom, and the target object can be the bedroom wall. The first image information can be a video frame captured by the user's camera of the bedroom. Therefore, each pixel in the first image information represents a group of points in the bedroom area captured by the camera. The first pose information includes the intrinsic and extrinsic parameters of the user's camera. Based on the first pose information, each pixel in the first image information can be projected into a group of points in world coordinates using the camera's intrinsic and extrinsic parameters. This group of points in world coordinates serves as the estimated points for each 3D spatial point. The distance information can then refer to the distance between each point in the aforementioned group of points in the world coordinate system and the wall. This distance information can be determined based on the first image information captured by the user and its corresponding first pose information.
[0088] In one embodiment, step 303 may specifically include: inputting the first image information and the first pose information into a preset prediction model, and outputting the distance information between each three-dimensional spatial point corresponding to each pixel in the first image information and the surface of the target object in the target spatial region.
[0089] In this embodiment, the distance information between each of the aforementioned three-dimensional spatial points in the target spatial region and the surface of the target object can be characterized by the TSDF value corresponding to each pixel in the first image information. At this time, the prediction model can be a pre-trained TSDF prediction network model (such as...). Figure 2B As shown, the first image information and the first pose information can be input into the TSDF prediction network model, and the distance information between each three-dimensional spatial point corresponding to each pixel and the surface of the target object in the target space region can be directly output. With the advantage of neural networks, it is convenient and fast.
[0090] In one embodiment, the step of training the prediction model includes: acquiring multi-channel features of a sample image and pose information of a preset frame image within the sample image; mapping the multi-channel features to a three-dimensional feature space based on the pose information of the preset frame image to obtain a three-dimensional feature map of the sample image; and training a preset neural network using the three-dimensional feature map to obtain the prediction model.
[0091] In this embodiment, the TSDF value output by the TSDF prediction network model represents the depth difference between a 3D point location and its nearest surface in a video frame captured by the camera. To learn the TSDF value, an FPN network is first used to learn from the input sample images and output the multi-channel features of the sample images. The preset frame image can be a single frame from the sample images. To improve the representativeness of the pose information, an intermediate frame image can be selected as the preset frame image. For example, if the sample images are a 4-second video with 60 frames per second, then 240 sample images are provided, and the 120th sample image can be selected as the preset frame image. Then, according to the pose information of the 120th sample image, the corresponding multi-channel features are mapped to the 3D feature space. Subsequently, 3D sparse convolution can be introduced for 3D feature learning. The ultimate goal is to enable the TSDF prediction network model to directly predict the TSDF value of the corresponding location point, and to optimize the prediction network model accordingly using regression loss (Korean L1 loss).
[0092] Step 304: Determine the depth information of each three-dimensional spatial point in the target spatial region based on the distance information and position estimation information.
[0093] In this step, since the position estimation information for each 3D spatial point is obtained by projecting the pose of a 2D image, the position estimation information often fails to provide accurate depth information in real-world scenarios. For example, if the 3D projection coordinates of each pixel in the first image information are (x, y, z), assuming x is the horizontal coordinate, y is the vertical coordinate, and z is the depth coordinate, then the estimated value of z is often inaccurate. The TSDF values of each pixel can compensate for this deficiency. The TSDF values of each pixel in step 303 can be used to supplement the position estimation information to determine the depth information of each corresponding 3D spatial point in the target spatial region. The resulting depth information accurately represents the position of each 3D spatial point in the target spatial region. Thus, accurate depth information can be obtained without additional input of a depth map, reducing the amount of data computation.
[0094] In one embodiment, step 304 may specifically include: determining the vector difference between the vector of position estimation information and the vector of distance information, and using the obtained vector difference as the depth information of each three-dimensional spatial point in the target spatial region.
[0095] In this embodiment, assuming that the estimated position information of a 3D spatial point corresponding to a pixel P in the first image information is a vector (x1, y1, z1), and the TSDF value corresponding to pixel P in step 303 is also a vector (x2, y2, z2), then the difference vector (x1-x2, y1-y2, z1-z2) obtained by subtracting vector (x2, y2, z2) from vector (x1, y1, z1) is the depth information of the 3D spatial point corresponding to pixel P in the target spatial region. Vector representation makes calculation simpler.
[0096] Step 305: Display the pixels corresponding to each 3D spatial point on the interactive interface in the first style according to the depth information.
[0097] In this step, the scanned points are the pixels corresponding to the three-dimensional spatial points that have been accurately captured in the first image information provided by the user. These three-dimensional spatial points have had their accurate depth information determined through the above process, so they are points with definite positions, and their corresponding pixels can be saved. The pixels corresponding to the three-dimensional spatial points that were not accurately captured are the unscanned points. In order to promptly remind the user which three-dimensional spatial points were missed, the pixels corresponding to the scanned three-dimensional spatial points can be displayed in the interactive interface using a preset first style. For example, the first style can be a highlight display. In this way, the scanned points and unscanned points in the first image information will be displayed differently on the interactive interface, making it easier for the user to see intuitively which areas were missed and improving the terminal's interactive performance.
[0098] In one embodiment, a 3D TSDF mesh (initialized to 0) can be used to store the results of scanned points. The depth information corresponding to the scanned points is rounded and stored in the 3D TSDF mesh. For scanned points, the corresponding position is assigned a value of 1, and for unscanned points, the corresponding position is assigned a value of 0. The TSDF mesh is actually a hash table. If the pixel corresponding to a 3D spatial point is in the 3D TSDF mesh, it means that it has been scanned and can be displayed in a bright color; otherwise, it means that it has not been scanned and unscanned points can be displayed in a dark color, providing a user-friendly prompt.
[0099] In one embodiment, scanned and unscanned points can be distinguished and displayed at corresponding positions on the image information. Assuming scanned points are displayed in bright colors and unscanned points in dark colors, scanned points can be directly displayed in bright colors on the image information, and unscanned points can be displayed in dark colors. This allows users to intuitively see which points have been scanned and which have not, facilitating the re-capture of unscanned points.
[0100] In one embodiment, the method may further include: processing individual pixels in the image information in parallel.
[0101] In this embodiment, the image information in steps 301 to 305 often includes multiple points. To improve processing speed, the processing of each pixel can be executed in parallel. For example, the indexing method and TSDF fusion method in steps 301 to 305 can be reimplemented using CUDA operators, and this can support operation on a GPU (graphics processing unit). Actual testing using CUDA operators shows a speedup of 28ms per image. This enables the image processing method to be implemented in real-time, supporting simultaneous 3D reconstruction and field indication tasks while reducing latency.
[0102] It should be noted that the execution order of steps 302 and 303 above is merely an example. In other embodiments, step 303 may be executed first, followed by step 302, or steps 302 and 303 may be executed simultaneously. This embodiment does not impose any limitations.
[0103] The aforementioned image data processing method uses an algorithm to determine whether the scene in the current video frame has been scanned, and accordingly colors it (unscanned areas are marked with a dark color, otherwise it is displayed normally). Users can use this coloring process as a prompt to determine whether the scene corresponding to the current location has been fully scanned, and to supplement any missed parts. A TSDF fusion approach is used, directly learning from a given depth map during the model training phase. This solution only requires depth map data during the training phase, and does not require depth maps during the testing phase, reducing the amount of data computation. In real-world scenarios, training the TSDF model is a fundamental requirement for 3D reconstruction tasks; therefore, the task of distinguishing between display and 3D reconstruction in the above embodiment can be integrated into a single end-to-end solution for unified resolution.
[0104] The aforementioned image data processing method can also be applied to scenarios of online purchase of interior decorations. For example, in the scenario of purchasing furniture, in order to show users how the selected furniture is placed in a virtual 3D scene, video frames of the home interior scene can be captured in advance to reconstruct the 3D model of the interior scene. During the shooting stage, the aforementioned image data processing method can be used to help users automatically determine whether the scene in the current video frame has been scanned and provide prompts, thereby improving the computing performance and interactive performance of the terminal and enhancing the user's furniture purchasing experience.
[0105] Please refer to Figure 4 This is an embodiment of an image data processing method according to this application. The method can be performed by... Figure 1 The electronic device 1 shown is used to perform this action and can be applied to... Figure 2A and Figure 2BIn the application scenario of the image data processing system shown, the goal is to determine the position of an object point in the target spatial region without the need for an additional depth map, and to distinguish between scanned and unscanned points to alert the user to missed areas. This eliminates the need for manual user judgment, improving the terminal's computing and interactive performance. This embodiment uses terminal 220 as the execution end. Compared to the previous embodiment, this embodiment also includes a process for re-shooting and providing retrieval prompts after the user's shooting is interrupted. This method includes the following steps:
[0106] Step 401: In response to the user's first operation on the target spatial region, acquire the first image information of the target spatial region and the first pose information corresponding to the first image information. See the detailed description of step 301 in the above embodiments.
[0107] Step 402: Based on the first pose information, perform 3D projection on each pixel in the first image information, and use the obtained 3D projected points as the position estimation information of each 3D spatial point in the target spatial region. See the detailed description of step 302 in the above embodiments.
[0108] Step 403: Input the first image information and the first pose information into the preset prediction model, and output the distance information between each three-dimensional spatial point corresponding to each pixel in the first image information and the surface of the target object in the target space region. See the detailed description of step 303 in the above embodiments.
[0109] Step 404: Determine the vector difference between the vector of the position estimation information and the vector of the distance information. The obtained vector difference is used as the depth information of each 3D spatial point in the target spatial region. See the detailed description of step 304 in the above embodiments.
[0110] Step 405: Display the pixels corresponding to each 3D spatial point on the interactive interface in the first style according to the depth information. See the detailed description of step 305 in the above embodiments.
[0111] Step 406: In response to the user's second operation on the target spatial region, acquire the second image information of the target spatial region and the second pose information corresponding to the second image information.
[0112] In this step, the second operation can be a scanning progress check of the target space area. For example, in step 401, a user opens their phone camera to take a picture of the bedroom, but is interrupted by other things. After the user finishes dealing with other things, they return to the bedroom to continue taking pictures of the bedroom for selection. However, the user doesn't remember where the interrupted shooting was. Therefore, a function to check the shooting progress can be triggered. For example, an icon for checking the shooting progress can be provided in the user interface. When the user presses the icon, the second operation is triggered, enabling the scanning progress check function. First, the phone camera can be opened to take a picture of the target space area, obtaining the second image information of the target space area and the second pose information corresponding to the second image information. That is, the user picks up the phone again to take a picture of the bedroom, continuously changing the direction and angle during the shooting process to ensure a more comprehensive picture. A detailed explanation of the second image information and its corresponding second pose information can be found in the description of the first image information and the first pose information in step 301 above.
[0113] Step 407: Based on the second image information and the second pose information, determine the depth information of the three-dimensional spatial points corresponding to each pixel point to be retrieved in the second image information in the target spatial region.
[0114] In this step, each pixel to be retrieved is a pixel captured in the second image information. Similar to steps 302 to 304, the depth information of the 3D spatial point corresponding to each pixel to be retrieved in the target spatial region can be determined. Multiple pixels to be retrieved can be processed in parallel. For example, based on the second pose information, each pixel to be retrieved in the second image information is 3D projected, and the resulting 3D projected points are used as the position estimation information of the 3D spatial point corresponding to each pixel to be retrieved in the target spatial region. The second image information and the second pose information are input into a preset prediction model, which outputs the distance information between the 3D spatial point corresponding to each pixel to be retrieved in the second image information and the surface of the target object in the target spatial region. The vector difference between the vector of the position estimation information and the vector of the distance information is used as the depth information of the 3D spatial point corresponding to each pixel to be retrieved in the target spatial region.
[0115] Step 408: Determine the scanning status of each pixel to be retrieved based on the depth information, and display the scanned points and unscanned points in the second image information on the interactive interface.
[0116] In this step, with the depth information of the three-dimensional spatial points corresponding to each pixel to be retrieved, the position information of the three-dimensional spatial points corresponding to each pixel to be retrieved in the target spatial region is determined. This allows the scanning status of each pixel to be retrieved to be determined, and the scanned and unscanned points in the second image information are displayed differently on the interactive interface. This allows users to intuitively know which areas were not captured, thereby assisting users in reshooting.
[0117] In one embodiment, step 408 may specifically include: determining, based on depth information, whether each pixel to be retrieved in the second image information is stored in the scanned point set; and displaying the first pixel to be retrieved in the second image information that is stored in the scanned point set on the interactive interface using a first style.
[0118] In this embodiment, the first point to be retrieved is the pixel corresponding to the three-dimensional spatial point that has been scanned in the second image information. The depth information of the three-dimensional spatial point corresponding to each pixel to be retrieved can be compared with the depth information of each point in the scanned point set saved in step 405. If there is a point in the scanned point set with the same depth information as the first point to be retrieved, it means that the three-dimensional spatial point corresponding to the first point to be retrieved has been scanned. The first point to be retrieved can be displayed on the interactive interface in a first style, such as highlighting the point at the position corresponding to the first point to be retrieved in the second image information.
[0119] In one embodiment, the step of determining whether each pixel to be retrieved in the second image information is stored in the scanned point set based on the depth information may specifically include: for each pixel to be retrieved in the second image information, searching in the scanned point set according to the corresponding depth information; if a point with the same depth information as the current pixel to be retrieved is found, it is determined that the current pixel to be retrieved is stored in the scanned point set; otherwise, it is determined that the current pixel to be retrieved is not stored in the scanned point set.
[0120] Specifically, the depth information of a pixel Q to be retrieved can be represented by coordinates. These coordinates are matched with the coordinates of the scanned points in the 3D TSDF mesh saved in step 405. If the position corresponding to the coordinates of the pixel Q to be retrieved in the 3D TSDF mesh is 1, it is determined that the pixel Q to be retrieved is saved in the set of scanned points, and it is determined that the pixel Q to be retrieved has been scanned. Otherwise, it is determined that the pixel Q to be retrieved is an unscanned point, and the user can be prompted that the 3D spatial point corresponding to point Q still needs to be scanned.
[0121] In one embodiment, step 408 may specifically include: displaying the second point to be retrieved in the second image information that is not saved in the scanned point set in a second style on the interactive interface, the second style being different from the first style.
[0122] In this embodiment, the second point to be retrieved refers to the pixel corresponding to an unscanned 3D spatial point in the second image information. If the coordinates of the pixel Q to be retrieved in the 3D TSDF grid are 0, then the pixel Q to be retrieved is determined to be the second point to be retrieved. This point has not been scanned and can be displayed on the interactive interface using the second style. The second style is to distinguish it from the first style. For example, the second style can be dark, mainly to prominently remind the user which points have been scanned and which points have not been scanned.
[0123] In one embodiment, the first style and the second style can take many forms, and any style that can be distinguished can be applied to this embodiment.
[0124] In one embodiment, when the entire area covered by the target space region is in the first pattern after multiple frames of scanning, it indicates that the scanning of the target space region is complete, and the user can be prompted and the scanning can be ended.
[0125] During this process, it can be determined whether the image captured by the user remains stationary for a period of time. For example, it can determine whether the captured image remains stationary within a preset number of frames. If so, a prompt can be issued, prompting the user to move and continue scanning. This avoids wasting resources due to user errors. The preset number of frames can be set based on practical experience, for example, 12 frames.
[0126] The aforementioned image data processing method, as the user scans, can update the scanning progress of the target space area with high accuracy and in real time to indicate the scan's completeness. Completed scanned areas of the target space area are represented by bright colors, while incomplete scans are represented by dark colors. Users can quickly achieve a realistic recreation of their home space, assisting them in selecting artwork online and efficiently changing paintings on the walls within the target space area. This significantly reduces decision-making costs and greatly enhances the certainty of purchasing artwork. By providing an immersive, real-time interactive service based on virtual reality or augmented reality, this method addresses the pain points of traditional art selection processes, such as high complexity and the inability to select artwork in real time. It offers users a multi-layered, highly realistic process experience, thereby better helping them choose suitable decorative paintings.
[0127] Please refer to Figure 5 This is an embodiment of an image data processing method according to this application. The method can be performed by... Figure 1 The electronic device 1 shown is used to perform this action and can be applied to... Figure 2A and Figure 2BIn the application scenario of the image data processing system shown, the goal is to determine the position of an object point in the target spatial region without the need for an additional depth map, and to distinguish between scanned and unscanned points to alert the user to missed areas. This eliminates the need for manual user judgment, improving the terminal's computing and interactive performance. This embodiment uses terminal 220 as the execution end, and the method includes the following steps:
[0128] Step 501: In response to the user's shooting operation on the indoor scene, acquire the first image information of the indoor scene and the first pose information corresponding to the first image information. See also Figure 2B The proposed scheme is shown in direction one.
[0129] Step 502: Based on the first image information and the first pose information, perform 3D reconstruction of the indoor scene. See [link / reference] Figure 2B The first option shown is the plan.
[0130] Step 503: Using the above... Figure 3 or Figure 4 The method processes each pixel in the first image information and distinguishes between scanned and unscanned points in the first image information. See the above embodiments for details. Figure 3 or Figure 4 Description of the embodiments.
[0131] The image data processing method described above integrates the scan point cues task and the 3D reconstruction of the target spatial region, allowing for a unified solution to both the 3D reconstruction and the target spatial region scan integrity cues algorithm in one go. The target spatial region scan integrity cues can be treated as a subtask of the target spatial region 3D reconstruction task. During the 3D reconstruction process, if the user needs to check the field scan integrity, they can do so in real time using the above method, resulting in simultaneous improvements in accuracy and efficiency.
[0132] Please refer to Figure 6 This is an image data processing apparatus 600 according to an embodiment of this application, which can be applied to... Figure 1 The electronic device 1 shown can be applied to Figure 2A and Figure 2B In the application scenario of the image data processing system shown, the goal is to determine the position of an object point in the target spatial region without the need for an additional depth map, and to distinguish between scanned and unscanned points to alert the user to missed areas. This eliminates the need for manual user judgment, improving the terminal's computing and interactive performance. The device includes: an acquisition module 601, a first determination module 602, a second determination module 603, a third determination module 604, and a display module 605. The underlying principles of each module are as follows:
[0133] The acquisition module 601 is used to respond to the user's first operation on the target spatial region, acquire the first image information of the target spatial region and the first pose information corresponding to the first image information, wherein the first image information includes at least: a group of pixels obtained by taking pictures of some three-dimensional spatial points in the target spatial region.
[0134] The first determining module 602 is used to determine the position estimation information of each three-dimensional spatial point corresponding to each pixel in the first image information in the target spatial region based on the first pose information and the first image information.
[0135] The second determining module 603 is used to determine the distance information between each three-dimensional spatial point and the surface of the target object in the target spatial region based on the first image information and the first pose information.
[0136] The third determining module 604 is used to determine the depth information of each three-dimensional spatial point in the target spatial region based on the distance information and the position estimation information.
[0137] Display module 605 is used to display the pixels corresponding to each three-dimensional spatial point on the interactive interface in a first style according to the depth information.
[0138] In one embodiment, the first determining module 602 is used to perform three-dimensional projection on each pixel in the first image information according to the first pose information, and the obtained three-dimensional projection points are used as position estimation information of each three-dimensional spatial point corresponding to each pixel in the target spatial region.
[0139] In one embodiment, the second determining module 603 is used to input the first image information and the first pose information into a preset prediction model, and output the distance information between each three-dimensional spatial point corresponding to each pixel in the first image information and the surface of the target object in the target space region.
[0140] In one embodiment, the system further includes a training module for training a prediction model, comprising: acquiring multi-channel features of a sample image and pose information of a preset frame image within the sample image; mapping the multi-channel features to a three-dimensional feature space based on the pose information of the preset frame image to obtain a three-dimensional feature map of the sample image; and training a preset neural network using the three-dimensional feature map to obtain the prediction model.
[0141] In one embodiment, the third determining module 604 is used to determine the vector difference between the vector of position estimation information and the vector of distance information, and the obtained vector difference is used as the depth information of each three-dimensional spatial point in the target spatial region.
[0142] In one embodiment, the system further includes: a retrieval module, configured to, in response to a second operation by a user on a target spatial region, acquire second image information of the target spatial region and second pose information corresponding to the second image information. Based on the second image information and the second pose information, the module determines the depth information of the three-dimensional spatial points corresponding to each pixel to be retrieved in the second image information within the target spatial region. Based on the depth information, the module determines the scanning state of each pixel to be retrieved and distinguishes between scanned and unscanned points in the second image information on the interactive interface.
[0143] In one embodiment, the retrieval module is used to determine, based on depth information, whether each pixel to be retrieved in the second image information is stored in the scanned point set. The first pixel to be retrieved in the second image information that is stored in the scanned point set is displayed on the interactive interface using a first style.
[0144] In one embodiment, the retrieval module is further configured to display the second retrieval point in the second image information that is not stored in the scanned point set in a second style on the interactive interface, the second style being different from the first style.
[0145] In one embodiment, the retrieval module is further configured to, for each pixel to be retrieved in the second image information, retrieve it in the scanned point set according to the corresponding depth information. If a point with the same depth information as the current pixel to be retrieved is found, it is determined that the current pixel to be retrieved is stored in the scanned point set; otherwise, it is determined that the current pixel to be retrieved is not stored in the scanned point set.
[0146] In one embodiment, distinguishing between scanned and unscanned points includes: displaying the scanned and unscanned points at corresponding positions on the image information.
[0147] In one embodiment, the method further includes: parallel processing of individual pixels in the image information.
[0148] For a detailed description of the image data processing device 600 described above, please refer to the description of the relevant method steps in the above embodiments. The implementation principle and technical effect are similar, and will not be repeated here in this embodiment.
[0149] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method of any of the foregoing embodiments.
[0150] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the foregoing embodiments.
[0151] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.
[0152] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.
[0153] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor. The memory may include high-speed RAM, and may also include non-volatile memory (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk, or optical disc, etc.
[0154] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.
[0155] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. Both the processor and the storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components in an electronic device or host device.
[0156] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0157] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0158] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0159] The collection, storage, use, processing, transmission, provision, and disclosure of user data and other information involved in the technical solution of this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0160] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. An image data processing method, characterized in that, The method includes: In response to a user’s first operation on a target spatial region, first image information of the target spatial region and first pose information corresponding to the first image information are obtained. The first image information includes at least: a group of pixels obtained by taking pictures of some three-dimensional spatial points in the target spatial region. The group of pixels is two-dimensional pixels. Based on the first pose information, each pixel in the first image information is projected into a three-dimensional space, and the resulting three-dimensional projection points are used as position estimation information of each three-dimensional space point in the target space region. The first image information and the first pose information are input into a preset prediction model, and the distance information between each three-dimensional spatial point corresponding to each pixel in the first image information and the surface of the target object in the target space region is output. The vector difference between the vector of the position estimation information and the vector of the distance information is determined, and the obtained vector difference is used as the depth information of each three-dimensional spatial point in the target spatial region; Each three-dimensional spatial point corresponding to the depth information is stored in the scanned point set, and the pixel points corresponding to each three-dimensional spatial point are displayed on the first image information in a first style. The prediction model is trained in the following manner: Acquire multi-channel features of the sample image and pose information of a preset frame image in the sample image, wherein the preset frame image is an intermediate frame image in the sample image frame; The multi-channel features are mapped to a three-dimensional feature space based on the pose information of the preset frame image to obtain the three-dimensional feature map of the sample image; The prediction model is obtained by training a preset neural network using the three-dimensional feature map.
2. The method according to claim 1, characterized in that, Also includes: In response to a second operation by the user on the target spatial region, second image information of the target spatial region and second pose information corresponding to the second image information are acquired; Based on the second image information and the second pose information, determine the depth information of the three-dimensional spatial points corresponding to each pixel point to be retrieved in the second image information in the target spatial region; The scanning status of each pixel to be retrieved is determined based on the depth information, and the scanned points and unscanned points in the second image information are displayed differently on the interactive interface.
3. The method according to claim 2, characterized in that, The step of determining the scanning status of each pixel to be retrieved in the second image information based on the depth information, and distinguishing between scanned and unscanned points in the second image information on the interactive interface, includes: Based on the depth information, determine whether each pixel to be retrieved in the second image information is stored in the scanned point set; The first retrieval point in the second image information that is stored in the scanned point set is displayed on the interactive interface using the first style.
4. The method according to claim 3, characterized in that, The step of determining the scanning status of each pixel to be retrieved in the second image information based on the depth information, and distinguishing between scanned and unscanned points in the second image information on the interactive interface, further includes: The second point to be retrieved, which is not saved in the scanned point set in the second image information, is displayed on the interactive interface in a second style, which is different from the first style.
5. The method according to claim 3, characterized in that, The step of determining whether each pixel to be retrieved in the second image information is stored in the scanned point set based on the depth information includes: For each pixel to be retrieved in the second image information, the corresponding depth information is used to search the scanned point set. If a point with the same depth information as the current pixel to be retrieved is found, it is determined that the current pixel to be retrieved is stored in the scanned point set; otherwise, it is determined that the current pixel to be retrieved is not stored in the scanned point set.
6. The method according to claim 1 or 2, characterized in that, Also includes: The individual pixels in the image information are processed in parallel.
7. An image data processing method, characterized in that, include: In response to the user's shooting operation of the indoor scene, the first image information of the indoor scene and the first pose information corresponding to the first image information are obtained; Based on the first image information and the first pose information, the indoor scene is reconstructed in three dimensions; Using the method described in any one of claims 1-6, each pixel in the first image information is processed, and the scanned and unscanned pixels in the first image information are displayed separately.
8. An electronic device, characterized in that, include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Mapping object instances using video data
CN112602116A
Object reconstruction method, device and apparatus and storage medium
CN112733579A
Data processing method and device, computer equipment and storage medium
CN114283243A