Data processing method and device, computer device, and storage medium

By generating a three-dimensional model with pixel information of the target space and using the virtual camera pose to generate panoramic roaming data, the problem of image distortion in panoramic roaming is solved, and more realistic and efficient space restoration and ranging are achieved.

CN114283243BActive Publication Date: 2025-10-10BEIJING JIUGUOCHUN INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111640797.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-10-10
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

Existing panoramic roaming technology is based on two-dimensional image stitching, which leads to large image distortion and cannot truly reflect the scene of the target space.

Method used

By generating a target three-dimensional model with pixel information in the target space and using the position of the virtual camera in the virtual scene to generate panoramic roaming data, direct mapping of pixel information is achieved and distortion is reduced.

Benefits of technology

The authenticity and flexibility of panoramic roaming data are improved, and the position and structure of objects in the target space can be restored more accurately. The ranging process is efficient and does not require complex calculations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114283243B_ABST
    Figure CN114283243B_ABST
Patent Text Reader

Abstract

The present disclosure provides a data processing method and device, computer equipment and a storage medium, wherein the method comprises: acquiring image data of a target space; generating a target three-dimensional model with pixel information of the target space based on the image data and a first pose of an image acquisition device in the target space when the image data is acquired; and in response to panoramic roaming being triggered, generating panoramic roaming data based on a second pose of a virtual camera in a virtual scene corresponding to the target space and the target three-dimensional model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of augmented reality (AR) technology, and in particular to a data processing method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the rapid development of AR technology, online panoramic roaming of certain places through panoramic views has become a relatively common means of sightseeing, such as AR house viewing and AR scenic spot tours. Current panoramic roaming technology is usually based on a camera capturing different two-dimensional images of a certain position point in the target space from multiple perspectives, and then stitching the two-dimensional images corresponding to the multiple perspectives into a panoramic image to obtain the panoramic roaming data corresponding to the position point; by generating panoramic roaming data corresponding to multiple position points, the panoramic roaming data of the target space is obtained. This method of generating panoramic roaming data has the problem of large image distortion and cannot reflect the real scene of the target space. Summary of the Invention

[0003] The embodiments of the present disclosure at least provide a data processing method, apparatus, computer equipment, and storage medium.

[0004] In a first aspect, an embodiment of the present disclosure provides a data processing method, comprising: acquiring image data of a target space; generating a target three-dimensional model of the target space with pixel information based on the image data and a first posture of an image acquisition device in the target space when acquiring the image data; and generating panoramic roaming data based on a second posture of a virtual camera in a virtual scene corresponding to the target space and the target three-dimensional model in response to panoramic roaming being triggered.

[0005] In this way, by generating a target three-dimensional model with pixel information of the target space and using the target three-dimensional model to generate panoramic roaming data of the target space, the panoramic roaming data can be generated by directly performing pixel mapping based on the pixel information of each vertex in the target three-dimensional model, thereby reducing the distortion of objects in the target space in the generated panoramic roaming data and restoring the target space more realistically.

[0006] In an optional embodiment, the generating of a target three-dimensional model of the target space with pixel information based on the image data and the first pose of the image acquisition device in the target space when acquiring the image data includes: performing three-dimensional reconstruction of the target space based on the image data and the first pose of the image acquisition device in the target space when acquiring the image data to obtain a three-dimensional model of the target space; the three-dimensional model includes: a plurality of vertices on the surface of the target object located in the target space, and position information of each vertex in the target space; performing pixel mapping on the vertices corresponding to each pixel point based on the pixel values ​​corresponding to each pixel point in the image data to obtain pixel information corresponding to each vertex in the three-dimensional model; obtaining the target three-dimensional model based on the pixel information corresponding to each vertex in the three-dimensional model.

[0007] In this way, based on the image data and the first pose of the image acquisition device in the target space when acquiring the image data, the target space is reconstructed in three dimensions to generate a three-dimensional model that can reflect the real spatial structure of the target space, the real spatial structure of each device contained in the target space, and the real position of each device in the target space. Based on the pixel information corresponding to each pixel point in the image data, the vertices corresponding to each pixel point are pixel mapped to obtain a target three-dimensional model of the target space carrying pixel information, which provides a more accurate data source for the subsequent generation of panoramic roaming data.

[0008] In an optional implementation, the target object includes at least one of the following: a building in the target space, and equipment deployed in the target space.

[0009] In an optional embodiment, in response to panoramic roaming being triggered, panoramic roaming data is generated based on the second posture of the virtual camera in the virtual scene corresponding to the target space and the target three-dimensional model, including: in response to panoramic roaming being triggered, based on the second posture of the virtual camera in the virtual scene corresponding to the target space, determining the viewing angle range corresponding to the virtual camera in the second posture; based on the target three-dimensional model, generating a projection image in the viewing angle range; based on the projection image in the viewing angle range, generating panoramic roaming data corresponding to the virtual camera in the second posture.

[0010] In this way, based on the target three-dimensional model and the second posture of the virtual camera in the virtual scene corresponding to the target space, a corresponding projection image is generated under the second posture, so as to realize three-dimensional perspective panoramic roaming of the target space based on the projection image, so that the panoramic roaming data can be roaming data corresponding to any position point in the target space, and therefore has greater flexibility.

[0011] In an optional embodiment, based on the target three-dimensional model, a projection image within the viewing angle range is generated, including: determining the projection plane corresponding to the virtual camera in the second posture; based on the posture of the projection plane in the target three-dimensional model, determining the target vertices from the target three-dimensional model; projecting the pixel information corresponding to the target vertices onto the projection plane to generate a projection image within the viewing angle range.

[0012] In an optional embodiment, after generating the panoramic roaming data, the method further includes: in response to a ranging operation on a first pixel point and a second pixel point in the panoramic roaming data, determining the distance between the first target vertex and the second target vertex based on a first depth value of the first target vertex corresponding to the first pixel point and a second depth value of the second target vertex corresponding to the second pixel point.

[0013] In this way, after the panoramic roaming data is generated, the first depth value of the first target vertex corresponding to the first pixel point and the second depth value of the second target vertex corresponding to the second pixel point can be determined in response to the ranging operation of the first pixel point and the second pixel point in the panoramic roaming data, and the first depth value and the second depth value can be used for ranging, so that the distance between different vertices in the target three-dimensional model can be accurately measured. At the same time, the ranging process does not require a complicated calculation process, has higher measurement efficiency, and achieves the effect of taking into account both measurement accuracy and measurement efficiency.

[0014] In an optional embodiment, the first pixel point and the second pixel point belong to the same frame of panoramic roaming data, and the method further includes: determining the depth information corresponding to each target vertex based on the second posture of the virtual camera in the virtual scene corresponding to the target space; rendering the depth information corresponding to each target vertex into a preset cache; and in response to the ranging operation of the first pixel point and the second pixel point in the panoramic roaming data, reading the first depth value corresponding to the first target vertex and the second depth value corresponding to the second target vertex from the preset cache.

[0015] In this way, when the first pixel point and the second pixel point belong to the same frame of panoramic roaming data, the depth information of each target vertex in the determined target three-dimensional model can be stored in a preset cache. In response to the ranging operation of the first pixel point and the second pixel point in the panoramic roaming data, the first depth value corresponding to the first target vertex and the second depth value corresponding to the second target vertex can be conveniently read from the preset cache, and the first depth value and the second depth value can be used to perform ranging, so that the distance between different vertices in the target three-dimensional model can be accurately measured. At the same time, the ranging process does not require a complex calculation process, has higher measurement efficiency, and achieves the effect of taking into account both measurement accuracy and measurement efficiency.

[0016] In an optional embodiment, the first pixel point belongs to a first frame of panoramic roaming data, the second pixel point belongs to a second frame of panoramic roaming data, and the first frame of panoramic roaming data and the second frame of panoramic roaming data are different panoramic roaming data. The method further includes: in response to a first pixel point selection operation on the first frame of panoramic roaming data, recording a first virtual camera pose corresponding to the first frame of panoramic roaming data and a first depth value of a first target vertex corresponding to the first pixel point; in response to a second pixel point selection operation on the second frame of panoramic roaming data, recording a second virtual camera pose corresponding to the second frame of panoramic roaming data and a second depth value of a second target vertex corresponding to the second pixel point.

[0017] In this way, by recording the first depth value of the first target vertex corresponding to the first pixel point and the first virtual camera pose, and the second depth value of the second target vertex corresponding to the second pixel point and the second virtual camera pose, it is convenient to accurately and efficiently determine the distance between the first target vertex and the second target vertex.

[0018] In an optional embodiment, in response to performing a first pixel selection operation on the first frame of panoramic roaming data, recording the first virtual camera pose corresponding to the first frame of panoramic roaming data and the first depth value of the first target vertex corresponding to the first pixel, includes: determining the depth information corresponding to each target vertex corresponding to the first frame of panoramic roaming data based on the first virtual camera pose of the virtual camera in the virtual scene corresponding to the target space when generating the first frame of panoramic roaming data; rendering the depth information corresponding to each target vertex into a preset cache; in response to performing a first pixel selection operation on the first frame of panoramic roaming data, reading the first depth value of the first target vertex corresponding to the first pixel from the preset cache, and recording the first virtual camera pose and the first depth value of the first target vertex.

[0019] In this way, the depth information of each target vertex in the determined target three-dimensional model is stored in a preset cache. In response to the selection operation of the first pixel point in the first frame of panoramic roaming data, the first depth value corresponding to the first target vertex can be conveniently read from the preset cache, thereby providing a data basis for the subsequent accurate measurement of the distance between the first target vertex and the second target vertex in the target three-dimensional model. At the same time, the distance measurement process does not require complex calculations, thereby improving the measurement efficiency, and further achieving the effect of taking into account both measurement accuracy and measurement efficiency.

[0020] In an optional embodiment, in response to performing a second pixel selection operation on the second frame of panoramic roaming data, recording the second virtual camera pose corresponding to the second frame of panoramic roaming data and the second depth value of the second target vertex corresponding to the second pixel, includes: determining the depth information corresponding to each target vertex corresponding to the second frame of panoramic roaming data based on the second virtual camera pose of the virtual camera in the virtual scene corresponding to the target space when generating the second frame of panoramic roaming data; rendering the depth information corresponding to each target vertex into a preset cache; in response to performing a second pixel selection operation on the second frame of panoramic roaming data, reading the second depth value of the second target vertex corresponding to the second pixel from the preset cache, and recording the second virtual camera pose and the second depth value of the second target vertex.

[0021] In this way, the depth information of each target vertex in the determined target three-dimensional model is stored in a preset cache. In response to the selection operation of the second pixel point in the second frame of panoramic roaming data, the second depth value corresponding to the second target vertex can be conveniently read from the preset cache, thereby providing a data basis for the subsequent accurate measurement of the distance between the first target vertex and the second target vertex in the target three-dimensional model. At the same time, the ranging process does not require complex calculations, thereby improving the measurement efficiency, and further achieving the effect of taking into account both measurement accuracy and measurement efficiency.

[0022] In a second aspect, an embodiment of the present disclosure also provides a data processing device, including: an acquisition module for acquiring image data of a target space; a first generation module for generating a target three-dimensional model of the target space with pixel information based on the image data and the first posture of the image acquisition device in the target space when acquiring the image data; a second generation module for generating panoramic roaming data based on the second posture of the virtual camera in the virtual scene corresponding to the target space and the target three-dimensional model in response to panoramic roaming being triggered.

[0023] In an optional embodiment, when the first generation module executes the generation of the target three-dimensional model of the target space with pixel information based on the image data and the first pose of the image acquisition device in the target space when acquiring the image data, it is specifically used to: perform three-dimensional reconstruction of the target space based on the image data and the first pose of the image acquisition device in the target space when acquiring the image data to obtain a three-dimensional model of the target space; the three-dimensional model includes: multiple vertices on the surface of the target object located in the target space, and position information of each vertex in the target space; based on the pixel values ​​corresponding to each pixel point in the image data, pixel mapping is performed on the vertices corresponding to each pixel point to obtain pixel information corresponding to each vertex in the three-dimensional model; based on the pixel information corresponding to each vertex in the three-dimensional model, the target three-dimensional model is obtained.

[0024] In an optional implementation, the target object includes at least one of the following: a building in the target space, and equipment deployed in the target space.

[0025] In an optional embodiment, when the second generation module generates panoramic roaming data in response to the panoramic roaming being triggered, based on the second posture of the virtual camera in the virtual scene corresponding to the target space and the target three-dimensional model, it is specifically used to: in response to the panoramic roaming being triggered, based on the second posture of the virtual camera in the virtual scene corresponding to the target space, determine the viewing angle range corresponding to the virtual camera in the second posture; based on the target three-dimensional model, generate a projection image in the viewing angle range; based on the projection image in the viewing angle range, generate the panoramic roaming data corresponding to the virtual camera in the second posture.

[0026] In an optional embodiment, when the second generation module generates a projection image within the viewing angle range based on the target three-dimensional model, it is specifically used to: determine the projection plane corresponding to the virtual camera in the second posture; determine the target vertices from the target three-dimensional model based on the posture of the projection plane in the target three-dimensional model; project the pixel information corresponding to the target vertices onto the projection plane to generate a projection image within the viewing angle range.

[0027] In an optional embodiment, after executing the generation of the panoramic roaming data, the second generation module is further used to: in response to the ranging operation of the first pixel point and the second pixel point in the panoramic roaming data, determine the distance between the first target vertex and the second target vertex based on the first depth value of the first target vertex corresponding to the first pixel point and the second depth value of the second target vertex corresponding to the second pixel point.

[0028] In an optional implementation, the first pixel point and the second pixel point belong to the same frame of panoramic roaming data, and the second generation module is further configured to: determine the depth information corresponding to each target vertex based on a second pose of the virtual camera in a virtual scene corresponding to the target space; render the depth information corresponding to each target vertex into a preset cache; and in response to a ranging operation on the first pixel point and the second pixel point in the panoramic roaming data, read the first depth value corresponding to the first target vertex and the second depth value corresponding to the second target vertex from the preset cache.

[0029] In an optional implementation, the first pixel point belongs to a first frame of panoramic roaming data, the second pixel point belongs to a second frame of panoramic roaming data, and the first frame of panoramic roaming data and the second frame of panoramic roaming data are different panoramic roaming data. The second generation module is further configured to: in response to a selection operation of a first pixel point on the first frame of panoramic roaming data, record a first virtual camera pose corresponding to the first frame of panoramic roaming data and a first depth value of a first target vertex corresponding to the first pixel point; and in response to a selection operation of a second pixel point on the second frame of panoramic roaming data, record a second virtual camera pose corresponding to the second frame of panoramic roaming data and a second depth value of a second target vertex corresponding to the second pixel point.

[0030] In an optional implementation, when performing the operation of recording the first virtual camera pose corresponding to the first frame of panoramic roaming data and the first depth value of the first target vertex corresponding to the first pixel point in response to the selection operation of the first pixel point on the first frame of panoramic roaming data, the second generation module is specifically configured to: determine the depth information corresponding to each target vertex based on a first virtual camera pose of the virtual camera in a virtual scene corresponding to the target space when the first frame of panoramic roaming data is generated; render the depth information corresponding to each target vertex into a preset cache; and in response to the selection operation of the first pixel point on the first frame of panoramic roaming data, read the first depth value of the first target vertex corresponding to the first pixel point from the preset cache.

[0031] In an optional embodiment, when the second generation module executes the recording of the second virtual camera pose corresponding to the second frame panoramic roaming data and the second depth value of the second target vertex corresponding to the second pixel point in response to the second pixel point selection operation on the second frame panoramic roaming data, it is specifically used to: determine the depth information corresponding to each target vertex corresponding to the second frame panoramic roaming data based on the second virtual camera pose of the virtual camera in the virtual scene corresponding to the target space when generating the second frame panoramic roaming data; render the depth information corresponding to each target vertex into a preset cache; in response to the second pixel point selection operation on the second frame panoramic roaming data, read the second depth value of the second target vertex corresponding to the second pixel point from the preset cache, and record the second virtual camera pose and the second depth value of the second target vertex.

[0032] In a third aspect, an optional implementation of the present disclosure further provides a computer device, a processor, and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is used to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the machine-readable instructions perform the steps of the above-mentioned first aspect, or any possible implementation of the first aspect.

[0033] In a fourth aspect, an optional implementation of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, it executes the steps of the above-mentioned first aspect or any possible implementation of the first aspect.

[0034] For a description of the effects of the above-mentioned data processing device, computer equipment, and computer-readable storage medium, please refer to the description of the above-mentioned data processing method, which will not be repeated here.

[0035] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without inventive effort.

[0037] Figure 1 A flow chart of a data processing method provided by an embodiment of the present disclosure is shown;

[0038] Figure 2 A flowchart showing a specific method of generating a three-dimensional dense model in the data processing method provided by an embodiment of the present disclosure is shown;

[0039] Figure 3 A schematic diagram of a data processing device provided by an embodiment of the present disclosure is shown;

[0040] Figure 4 A schematic diagram of a computer device provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0041] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The components of the embodiments of the present disclosure generally described and shown here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the present disclosure for protection, but merely represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present disclosure.

[0042] Research has found that it is currently possible to capture space through a panoramic camera device with at least two cameras, where the sum of the viewing angles of all camera lenses forms a spherical viewing angle greater than or equal to 360 degrees. The captured images are then transmitted to an image processing terminal, and the image processing software is used to modify the junction of the images captured by different cameras so that the images captured by different cameras are smoothly combined to generate a 360-degree panoramic image, and panoramic roaming is achieved based on the 360-degree panoramic image. However, since the images captured by the cameras are two-dimensional images, objects will be distorted to a certain extent when forming the panoramic image, resulting in the panoramic image being unable to reflect the real scene of the space.

[0043] Based on the above research, the present disclosure provides a data processing method, device, computer equipment and storage medium, which generates a target three-dimensional model of the target space with pixel information, and uses the target three-dimensional model to generate panoramic roaming data of the target space, so that the panoramic roaming data can be generated by directly performing pixel mapping based on the pixel information of each vertex in the target three-dimensional model, thereby reducing the distortion of objects in the target space in the generated panoramic roaming data and restoring the target space more realistically.

[0044] The defects of the existing solutions and the solutions proposed in this disclosure are the results obtained by the inventor after practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed in this disclosure for the above problems should be the contributions made by the inventor to this disclosure during the disclosure process.

[0045] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0046] To facilitate understanding of this embodiment, a data processing method disclosed in an embodiment of the present disclosure is first introduced in detail. The execution subject of the data processing method provided in the embodiment of the present disclosure is generally a computer device with certain computing capabilities. The computer device includes, for example, a terminal device or a server or other processing device. The terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementations, the data processing method can be implemented by a processor calling computer-readable instructions stored in a memory.

[0047] See also Figure 1 FIG. 1 is a flow chart of a data processing method provided by an embodiment of the present disclosure, wherein the method includes steps S101 to S103, wherein:

[0048] S101: Acquire image data of a target space.

[0049] Among them, the target space may include but is not limited to indoor physical spaces such as computer rooms (or sites), factories, and exhibition halls. Taking the target space including a computer room as an example, computer equipment, data storage equipment, and signal receiving equipment, etc. may be stored therein; taking the target space including a factory as an example, production equipment, loading and unloading equipment, transportation equipment, etc. may be present therein; in addition, the target space may also be an outdoor physical space, for example, the surrounding environment of a tower used for communication or for transmitting electricity is detected to prevent the vegetation around the tower from affecting the normal use of the tower during its growth process. Therefore, the space where the tower and the surrounding environment are located can be used as the target space; the image data includes but is not limited to images or videos obtained by using an image acquisition device to acquire images of the target space; wherein, the video may include panoramic video, and the image acquisition device may include but is not limited to at least one of a mobile phone, a camera, a camcorder, a panoramic camera, an unmanned aerial vehicle, etc.

[0050] For example, when the image data includes video, a robot equipped with an image acquisition device can be controlled to walk in the target space to obtain a video corresponding to the target space; or, surveyors or other staff can use image acquisition devices to capture images of the target space to obtain a video corresponding to the target space; or, a drone equipped with an image acquisition device can be controlled to fly in the target space to capture a video of the target space.

[0051] Here, when capturing images of the target space, in order to ensure the modeling quality of the target space, the image capture device can be controlled to capture images of the target space at different postures to form a video corresponding to the target space.

[0052] In particular, because the video frame images obtained by the image acquisition device through image acquisition of the target space are used for data processing, such as for reconstructing a three-dimensional model, it is necessary to determine the first position of the image acquisition device in the target space when acquiring the image data. In this case, for example, before using the image acquisition device to capture images of the target space, the gyroscope of the image acquisition device can be calibrated to determine the first position of the image acquisition device in the target space. For example, the optical axis of the image acquisition device can be adjusted to be parallel to the ground surface of the target space.

[0053] After the gyroscope of the image acquisition device is calibrated, the image acquisition device can be selected in a video mode to perform image acquisition and obtain a video corresponding to the target space.

[0054] Following S101 above, the data processing method provided by the embodiment of the present disclosure further includes:

[0055] S102: Generate a three-dimensional target model with pixel information in the target space based on the image data and the first pose of the image acquisition device in the target space when acquiring the image data.

[0056] In a specific implementation, a target three-dimensional model with pixel information of the target space can be generated in the following manner: based on the image data and the first pose of the image acquisition device in the target space when acquiring the image data, the target space is three-dimensionally reconstructed to obtain a three-dimensional model of the target space; based on the pixel values ​​corresponding to each pixel point in the image data, pixel mapping is performed on the vertices corresponding to each pixel point to obtain pixel information corresponding to each vertex in the three-dimensional model; based on the pixel information corresponding to each vertex in the three-dimensional model, the target three-dimensional model is obtained.

[0057] Among them, the three-dimensional model includes: multiple vertices on the surface of the target object located in the target space, and the position information of each vertex in the target space; here, the target object in the target space may include but is not limited to at least one of the following: buildings in the target space, and equipment deployed in the target space; taking the target space including the computer room as an example, the buildings in the target space may include but are not limited to at least one of the computer room walls, computer room ceilings, computer room floors, computer room columns, etc.; the equipment deployed in the target space may include but is not limited to: pole towers and outdoor cabinets installed on the roof of the computer room, and wiring racks connected to the pole towers, and indoor cabinets placed on the floor of the computer room, etc.

[0058] For example, when performing three-dimensional reconstruction of the target space based on the film and television data and the first pose of the image acquisition device in the target space when acquiring the film and television data to obtain a three-dimensional model of the target space, at least one of the following methods A1 to A2 may be used, but is not limited to:

[0059] A1. The image acquisition device only undertakes the task of image acquisition and relies on the network connection to transmit the acquired image data and the first pose in the target space when acquiring the image data to the data processing device, so that the data processing device can establish a three-dimensional model of the target space.

[0060] Among them, the network connection that can be relied upon may include, but is not limited to, a fiber optic Ethernet adapter, mobile communication technology (such as the fourth generation mobile communication technology (4G) or the fifth generation mobile communication technology (5G)), and wireless fidelity communication (Wireless Fidelity, Wi-Fi); the data processing equipment may include, but is not limited to, the computer equipment described above.

[0061] When the data processing device processes the image data, for example, it can perform a three-dimensional dense reconstruction of the target space based on the image data and the first pose of the image acquisition device when acquiring the image data (i.e., the pose of the image acquisition device in the target space), thereby obtaining a three-dimensional dense model of the target space. Here, the three-dimensional dense model includes a plurality of facets and vertices corresponding to the plurality of facets; any facet has at least one common vertex with at least one other facet; the facets provided in the embodiments of the present disclosure may include, but are not limited to, at least one of triangular facets and quadrilateral facets, etc., without specific limitation.

[0062] Here, when obtaining the first position of the image acquisition device in the target space when collecting image data, for example, relevant data of the inertial measurement unit (IMU) of the image acquisition device when collecting image data can be obtained. Exemplarily, the inertial measurement unit IMU of the image acquisition device may include, for example, three single-axis accelerometers and three single-axis gyroscopes. The accelerometer can detect the acceleration of the image acquisition device when collecting image data, and the gyroscope can detect the angular velocity of the image acquisition device when collecting image data. In this way, by collecting relevant data of the inertial measurement unit IMU in the image acquisition device, the first position of the image acquisition device when collecting image data can be accurately determined.

[0063] For example, when the image acquisition device collects image data, a three-dimensional dense model covering the target object can be gradually generated as the image acquisition device gradually moves to collect image data; or, after the image acquisition device finishes collecting image data, the obtained complete image data is used to generate a three-dimensional dense model corresponding to the target object.

[0064] In another embodiment of the present disclosure, a three-dimensional model of the target space can be generated based on a panoramic video obtained by capturing images of the target space using a panoramic camera. The panoramic cameras are two fisheye cameras positioned in front and back of a scanner. The fisheye cameras are positioned in a preset position on the scanner to capture a complete panoramic video of the target space.

[0065] See also Figure 2 FIG. 1 is a flowchart of a specific method for generating a three-dimensional dense model by using a panoramic video obtained by a data processing device provided by an embodiment of the present disclosure to capture an image of a target object using a panoramic camera, wherein:

[0066] S201: A data processing device obtains two panoramic videos that are captured in real time by two fisheye cameras in front and behind the scanner and are synchronized in time.

[0067] The two panoramic videos each include multiple video frames. Since the two fisheye cameras capture the two panoramic videos in real time and in synchronization with each other, the timestamps of the multiple video frames in the two panoramic videos are corresponding.

[0068] Additionally, the timestamp accuracy and the acquisition frequency of the panoramic video frames can be determined based on the specific instrument parameters of the two fisheye cameras. For example, the timestamp accuracy of the panoramic video frames can be set to nanoseconds, and the acquisition frequency of the panoramic video frames can be set to be no less than 30 Hz.

[0069] S202: The data processing device determines relevant data of the inertial measurement units (IMUs) of the two fisheye cameras when respectively acquiring panoramic videos.

[0070] Taking either of the two fisheye cameras as an example, when capturing video frames in a panoramic video, the IMU data between two adjacent video frames can be observed and acquired, along with the timestamp of when the data was acquired. Specifically, a corresponding scanner coordinate system (e.g., consisting of an X-axis, a Y-axis, and a Z-axis) can be determined for the fisheye camera to determine IMU data in the scanner coordinate system, such as acceleration and angular velocity along the X, Y, and Z axes of the scanner coordinate system.

[0071] In addition, the timestamp for acquiring the relevant data of the inertial measurement unit (IMU) can be determined based on the specific device parameters of the two fisheye cameras. For example, the observation frequency for acquiring the relevant data of the inertial measurement unit (IMU) can be determined to be no less than 400 Hz.

[0072] S203: The data processing device determines the positions of the two fisheye cameras in the world coordinate system based on the relevant data of the inertial measurement unit IMU.

[0073] Specifically, since the coordinate system conversion relationship between the scanner coordinate system and the world coordinate system can be determined, after obtaining the relevant data of the inertial measurement unit IMU, the position and posture of the two fisheye cameras in the world coordinate system can be determined based on the coordinate system conversion relationship. For example, it can be expressed as a 6-degree of freedom (6DOF) position. Specifically, based on the coordinate system conversion relationship between the scanner coordinate system and the world coordinate system, the position and posture of the two fisheye cameras in the world coordinate system are determined using the existing coordinate system conversion method, which will not be repeated here.

[0074] Regarding the above S201 to S203, since the video frame images in the panoramic video are all panoramic images, the processing steps of image processing, key point extraction, key point tracking, and establishing the correlation relationship between key points can be used to accurately solve the 6DOF pose of the image acquisition device, that is, real-time 6DOF pose acquisition and calculation of the image acquisition device is realized; and the coordinates of the dense point cloud points in the target space can also be obtained.

[0075] Among them, when processing the video frame images in the panoramic video, the key frame images can also be determined in the corresponding multiple video frame images in the panoramic video, so as to reduce the amount of calculation and improve efficiency while ensuring a sufficient amount of processing data in the process of three-dimensional dense reconstruction.

[0076] Specifically, the method for determining the key frame image from the panoramic video may be, for example, but not limited to, at least one of the following methods B1 to B4:

[0077] B1. Using an alternate frame sampling method, extract at least one frame of video image from the panoramic video as a key frame image.

[0078] B2. Extracting a preset number of video frame images at a preset time, and extracting at least one video frame image from the panoramic video as a key frame image.

[0079] The extraction of a preset number of video frame images within a preset time may include, but is not limited to, two frames per second.

[0080] B3. Use image processing algorithms, image analysis algorithms, natural language processing (NLP) and other technologies to identify the content of each frame of video in the panoramic video, determine the semantic information corresponding to each frame of video, and extract the video frame image containing the target object as the key frame image based on the semantic information corresponding to each frame of video.

[0081] B4. In response to the selection of the video frame image in the panoramic video, determine the key frame image in the panoramic video.

[0082] In a specific implementation, a panoramic video of the target space can be displayed to the user, and when the panoramic video is displayed, in response to the user's selection operation on part of the video frames, the selected part of the video frames can be used as key frame images in the panoramic video.

[0083] For example, when displaying a panoramic video to a user, a prompt message for selecting a key frame image can be displayed to the user. Specifically, for example, in response to a user's long press, double-click, or other specific operation, a video frame image in the panoramic video can be selected, and the selected video frame image can be used as a key frame image. In addition, a prompt message can be displayed, such as a message containing the text "Please long press to select this video frame image", and when the user long presses any video frame image in the panoramic video, the video frame image is used as a key frame image.

[0084] After determining the key frame images in the panoramic video, the key frame image map can be stored in the background so that after controlling the image acquisition device to return to the acquired position, the two video frame images at that position can be compared to perform loop detection on the image acquisition device, thereby correcting the accumulated positioning error of the image acquisition device under long-term and long-distance operations.

[0085] S204, the data processing device takes the key frame image in the panoramic video obtained by the fisheye camera respectively and the pose of the fisheye camera as input data of the real-time dense reconstruction algorithm for processing.

[0086] For example, for any panoramic video obtained by the fisheye camera, after determining the new key frame image in the panoramic video by S201-S203, all the key frame images obtained at present and the pose of the fisheye camera corresponding to the new key frame image are taken as input data of the real-time dense reconstruction algorithm.

[0087] Among them, since the pose of the fisheye camera corresponding to the key frame image has been input to the real-time dense reconstruction algorithm as input data when the key frame image has been transmitted, the input of the new key frame image can no longer be repeated.

[0088] S205, the data processing device processes the input data by using the real-time dense reconstruction algorithm to obtain a three-dimensional dense model corresponding to the target space.

[0089] For example, the obtained three-dimensional dense model may include but is not limited to: a plurality of surface patches on the surface of the target object, and the vertices corresponding to each surface patch, and the position information of each vertex in the target space.

[0090] In the above S204-S205, when the real-time dense reconstruction algorithm is used, the dense stereo matching technology can be used to estimate the dense depth map corresponding to the key frame image, and the dense depth map can be fused into a three-dimensional dense model by using the pose of the corresponding fisheye camera, so as to generate a three-dimensional model after the target space is collected.

[0091] Among them, the dense depth map is also called distance image, which is different from the pixel point in the gray image which stores the brightness value, and the pixel point stores the distance between the point and the image acquisition device, that is, the depth value; since the size of the depth value is only related to the distance, and has nothing to do with the environment, light, direction and other factors, the dense depth map can truly and accurately reflect the geometric depth information of the target space, so that a three-dimensional dense model reflecting the real target space can be generated based on the dense depth map; in addition, considering the limitation of device resolution, the dense depth map can be denoised or repaired for image enhancement processing to provide high-quality dense depth image for three-dimensional reconstruction.

[0092] In one possible scenario, for a processed keyframe image, using the pose of the image acquisition device corresponding to the keyframe image and the pose of the image acquisition device corresponding to the adjacent new keyframe image, it is possible to determine whether the pose of the image acquisition device has been adjusted when acquiring the target space. If the pose has not been adjusted, then real-time 3D dense reconstruction of the target space is continued to obtain a 3D dense model; if the pose has been adjusted, then the dense depth map is adjusted accordingly according to the pose adjustment, so as to perform real-time 3D dense reconstruction of the target space based on the adjusted dense depth map, thereby obtaining an accurate 3D dense model.

[0093] A2. The image acquisition device has computing power to process image data. After acquiring the image data, the device uses its own computing power to process the image data to obtain a three-dimensional model corresponding to the target space.

[0094] Here, the specific manner in which the image acquisition device generates a three-dimensional model of the target space based on the image data can refer to the description of A1 above, and the repeated parts will not be repeated.

[0095] After performing three-dimensional reconstruction of the target space based on at least one of A1 to A2 above to obtain a three-dimensional model of the target space, the vertices corresponding to each pixel point in the film and television data in the three-dimensional model can be determined, and based on the pixel values ​​corresponding to each pixel point in the film and television data, pixel mapping is performed on the vertices corresponding to each pixel point to obtain pixel information corresponding to each vertex in the three-dimensional model; based on the pixel information corresponding to each vertex in the three-dimensional model, a target three-dimensional model of the target space with pixel information is obtained.

[0096] Following S102 above, the data processing method provided by the embodiment of the present disclosure further includes:

[0097] S103 : In response to the panoramic roaming being triggered, generate panoramic roaming data based on the second position of the virtual camera in the virtual scene corresponding to the target space and the target three-dimensional model.

[0098] In the embodiments of the present disclosure, in order to solve the problem in the prior art that objects in the panoramic picture generated by using a two-dimensional image of the target space captured by an image acquisition device may be deformed, resulting in the inability to roam in the real scene corresponding to the target space based on the panoramic picture, a three-dimensional target model with pixel information that can reflect the real target space can be used to generate panoramic roaming data of the target space.

[0099] In a specific implementation, in response to the panoramic roaming being triggered, the viewing angle range corresponding to the virtual camera in the second posture is determined based on the second posture of the virtual camera in the virtual scene corresponding to the target space; based on the target three-dimensional model, a projection image in the viewing angle range is generated; based on the projection image in the viewing angle range, the panoramic roaming data corresponding to the virtual camera in the second posture is generated.

[0100] Among them, the second posture of the virtual camera in the virtual scene corresponding to the target space may include position and posture. The second posture can be, for example, but not limited to: the first posture when the image acquisition device acquires image data of the target space; the viewing angle range corresponding to the virtual camera in the second posture may include but is not limited to: placing the shooting origin of the virtual camera at the position corresponding to the second posture, and determining the shooting range of the virtual camera based on the posture of the virtual camera in the second posture.

[0101] Specifically, based on the second posture of the virtual camera in the virtual scene corresponding to the target space, the projection plane corresponding to the virtual camera in the second posture is determined; based on the posture of the projection plane in the target three-dimensional model, the target vertex is determined from the target three-dimensional model; the pixel information corresponding to the target vertex is projected onto the projection plane to generate a projection image within the viewing angle range.

[0102] Exemplarily, the position in the second posture can be used as the shooting origin of the virtual camera, and the shooting range (i.e., the viewing angle range) of the virtual camera in the second posture is determined based on the position information of the shooting origin and the focal length of the virtual camera; based on the second posture and the shooting range of the virtual camera in the second posture, the projection plane corresponding to the virtual camera in the second posture is determined; in the target three-dimensional model, the target vertex that matches the posture of the projection plane is determined as the vertex that can be projected onto the projection plane; the pixel information carried by each determined vertex is projected onto the projection plane, thereby generating a projection image under the viewing angle range; the virtual camera is controlled to move in the virtual scene corresponding to the target space, and the second posture of the virtual camera in the virtual scene corresponding to the target space is continuously changed, so as to obtain projection images under different viewing angle ranges and obtain panoramic roaming data.

[0103] In a specific implementation, after generating the panoramic roaming data, since each vertex in the target three-dimensional model used when generating the panoramic roaming data carries depth information, the distance between the vertices corresponding to any two pixel points in the panoramic roaming data can be determined. Specifically: in response to the ranging operation of the first pixel point and the second pixel point in the panoramic roaming data, the distance between the first target vertex and the second target vertex is determined based on the first depth value of the first target vertex corresponding to the first pixel point and the second depth value of the second target vertex corresponding to the second pixel point.

[0104] In many scenarios, it's necessary to determine the distance between real objects and the corresponding three-dimensional models using a three-dimensional model constructed of real physical space. Current distance measurement methods include the following: First, a bounding box is generated for each corresponding three-dimensional model of a real object; the corresponding three-dimensional model of the real object is located within the bounding box; and the distance between the bounding boxes is determined to determine the distance between the real objects. While this method is relatively simple to measure, there may actually be a large gap between the bounding box and the three-dimensional model corresponding to the real object, resulting in large errors in the distance determined based on the bounding box. Second, a three-dimensional model is constructed using point clouds or grids. For example, using a grid requires traversing each grid in the three-dimensional model, determining the distance corresponding to each grid, and then determining the distance between objects based on the distances corresponding to each grid. While this method provides high accuracy, the number of grids that make up the three-dimensional model is large, and traversing each grid in the three-dimensional model sequentially takes a long time. This results in a large amount of computation and processing time required to determine the distance, resulting in low measurement efficiency and a high consumption of computing resources. Therefore, current distance measurement methods suffer from the problem of being unable to balance measurement accuracy and efficiency.

[0105] In the embodiment of the present disclosure, the first depth value of the first pixel point and the second depth value of the second pixel point are utilized to obtain the distance between the first target vertex corresponding to the first pixel point and the second target vertex corresponding to the second pixel point. Compared with the ranging method in the related art, it has higher measurement accuracy and higher measurement efficiency, thereby achieving the effect of taking into account both measurement accuracy and measurement efficiency.

[0106] For example, the embodiment of the present disclosure may adopt, but is not limited to, at least one of the following C1 to C2 to determine the distance between vertices corresponding to any two pixels in the panoramic roaming data:

[0107] C1. When the first pixel point and the second pixel point belong to the same frame of panoramic roaming data, the depth information corresponding to each target vertex can be determined based on the second posture of the virtual camera in the virtual scene corresponding to the target space; the depth information corresponding to each target vertex is rendered into a preset cache; in response to the ranging operation of the first pixel point and the second pixel point in the panoramic roaming data, the first depth value corresponding to the first target vertex and the second depth value corresponding to the second target vertex are read from the preset cache; based on the first depth value and the second depth value, the distance between the first target vertex and the second target vertex is determined.

[0108] The preset cache, for example, is located in the graphics card's corresponding cache and includes: a cache space for storing pixel values ​​and depth values ​​for each pixel, including the Gbuffer; the Gbuffer is where the graphics card stores the depth values ​​of pixels. For the geometry rendering stage, we first need to initialize a framebuffer object, also known as the gBuffer, which contains multiple color buffers and a separate depth renderbuffer object.

[0109] For example, when generating panoramic roaming data of a virtual camera in a second posture based on the second posture of the virtual camera in the virtual scene corresponding to the target space and the target three-dimensional model in response to a display trigger operation of the panoramic roaming data, the depth information corresponding to each target vertex can be determined based on the second posture of the virtual camera in the virtual scene corresponding to the target space; the projection relationship information between the virtual camera and the target three-dimensional model can be determined based on the second posture in the virtual scene corresponding to the target space and the posture of the target three-dimensional model; the projection position of each target vertex in the projected image can be determined based on the projection relationship information; the cache position of each target vertex in the preset cache can be determined based on the projection position; and the depth information corresponding to the target vertex can be rendered into the preset cache based on the cache position of each target vertex in the preset cache.

[0110] The projection relationship information between the virtual camera and the target three-dimensional model refers to the mapping relationship when each vertex in the target three-dimensional model is projected onto the projection plane corresponding to the virtual camera.

[0111] After rendering the depth information corresponding to the target vertex into a preset cache, in response to the ranging operation of the first pixel and the second pixel in the projected image (i.e., the panoramic roaming data of the virtual camera in the second posture), based on the first position of the first pixel in the projected image, a first storage position of the first depth value of the first target vertex corresponding to the first pixel is determined from the preset cache; based on the first storage position, the first depth value is read from the preset cache; and, based on the second position of the second pixel in the projected image, a second storage position of the second depth value of the second target vertex corresponding to the second pixel is determined from the preset cache; based on the second storage position, the second depth value is read from the preset cache.

[0112] After reading the first depth value and the second depth value from the preset cache, a first distance of the first target vertex and the second target vertex in depth can be determined based on the first depth value and the second depth value; and a second distance of the first target vertex and the second target vertex after being projected to a projection plane on which the projection image is located can be determined based on a distance of the first pixel point and the second pixel point on the projection plane and a camera projection principle; a third distance of the first target vertex and the second target vertex in a model coordinate system can be determined based on the first distance and the second distance; and a distance between the first target vertex and the second target vertex can be determined based on a proportional relationship between the model coordinate system and a real space and the third distance.

[0113] Specifically, after the first distance and the second distance are calculated, a third distance between the first target vertex and the second target vertex in the three-dimensional model can be determined based on the Pythagorean theorem and the first distance and the second distance. For example, a formula for calculating the third distance of the first target vertex and the second target vertex in the model coordinate system can be shown in Formula 1:

[0114]

[0115] wherein, l1 represents the third distance of the first target vertex and the second target vertex in the model coordinate system; h1 represents the first depth value of the first target vertex; h2 represents the second depth value of the second target vertex; |h1-h2| represents the first distance of the first target vertex and the second target vertex in depth; and c represents the second distance of the first target vertex and the second target vertex on the projection plane on which the projection image is located.

[0116] After the third distance of the first target vertex and the second target vertex in the model coordinate system is determined, the distance between the first target vertex and the second target vertex can be determined based on the proportional relationship between the model coordinate system and the real space and the third distance. The proportional relationship between the model coordinate system and the real space can be set according to actual needs, which is not limited here; and the distance includes the distance of the first target vertex and the second target vertex in the target space.

[0117] For example, if the proportional relationship between the model coordinate system and the real space includes 1:10, after the third distance l1 of the first target vertex and the second target vertex in the model coordinate system is determined, the distance l2 between the first target vertex and the second target vertex can be calculated according to l2=10l1.

[0118] C2. When the first pixel point belongs to the first frame of panoramic roaming data, the second pixel point belongs to the second frame of panoramic roaming data, and the first frame of panoramic roaming data and the second frame of panoramic roaming data are different panoramic roaming data, in response to a selection operation of the first pixel point on the first frame of panoramic roaming data, the first virtual camera pose corresponding to the first frame of panoramic roaming data and the first depth value of the first target vertex corresponding to the first pixel point can be recorded; in response to a selection operation of the second pixel point on the second frame of panoramic roaming data, the second virtual camera pose corresponding to the second frame of panoramic roaming data and the second depth value of the second target vertex corresponding to the second pixel point can be recorded; and based on the first depth value and the second depth value, the distance between the first target vertex and the second target vertex can be determined.

[0119] In a specific implementation, the first virtual camera pose corresponding to the first frame of panoramic roaming data and the first depth value of the first target vertex corresponding to the first pixel point can be recorded in the following manner: based on the first virtual camera pose of the virtual camera in the virtual scene corresponding to the target space when generating the first frame of panoramic roaming data, the depth information corresponding to each target vertex corresponding to the first frame of panoramic roaming data is determined; the depth information corresponding to each target vertex is rendered into a preset cache; in response to the selection operation of the first pixel point of the first frame of panoramic roaming data, the first depth value of the first target vertex corresponding to the first pixel point is read from the preset cache, and the first virtual camera pose and the first depth value of the first target vertex are recorded.

[0120] Exemplarily, in response to a display trigger operation for the first frame of panoramic roaming data, the depth information corresponding to each target vertex corresponding to the first frame of panoramic roaming data can be determined based on the first virtual camera pose of the virtual camera in the virtual scene corresponding to the target space when the first frame of panoramic roaming data is generated; the projection relationship information between the virtual camera and the target three-dimensional model can be determined based on the first virtual camera pose and the pose of the target three-dimensional model; the projection position of each target vertex in the first frame of panoramic roaming data can be determined based on the projection relationship information; the cache position of each target vertex in the preset cache can be determined based on the projection position; and the depth information corresponding to the target vertex can be rendered into the preset cache based on the cache position of each target vertex in the preset cache.

[0121] After rendering the depth information corresponding to each target vertex corresponding to the first frame of panoramic roaming data into a preset cache, in response to a first pixel selection operation on the first frame of panoramic roaming data, based on the first position of the first pixel in the first frame of panoramic roaming data, a first storage location of a first depth value of a first target vertex corresponding to the first pixel is determined from the preset cache; based on the first storage location, the first depth value is read from the preset cache; and the first virtual camera pose and the first depth value are recorded.

[0122] After recording the first virtual camera pose and the first depth value, the pose of the virtual camera can be adjusted to generate second frame panoramic roaming data, and a second pixel point is selected in the second frame panoramic roaming data to measure the distance between the first target vertex corresponding to the first pixel point and the second target vertex corresponding to the second pixel point.

[0123] In a specific implementation, the second virtual camera pose corresponding to the second frame panoramic roaming data and the second depth value of the second target vertex corresponding to the second pixel point can be recorded in the following manner: based on the second virtual camera pose of the virtual camera in the virtual scene corresponding to the target space when the second frame panoramic roaming data is generated, the depth information corresponding to each target vertex is determined respectively; the depth information corresponding to each target vertex is rendered to the preset cache respectively; in response to the selection operation of the second pixel point on the second frame panoramic roaming data, the second depth value of the second target vertex corresponding to the second pixel point is read from the preset cache, and the second virtual camera pose and the second depth value of the second target vertex are recorded.

[0124] Here, the specific implementation of storing the depth information corresponding to each target vertex corresponding to the second frame panoramic roaming data into the preset cache and reading the second depth value of the second target vertex from the preset cache is similar to the specific implementation of storing the depth information corresponding to each target vertex corresponding to the first frame panoramic roaming data into the preset cache and reading the first depth value of the first target vertex from the preset cache in the embodiments of the present disclosure, and the repeated parts will not be described again.

[0125] After recording the first depth value of the first target vertex corresponding to the first pixel point in the first frame of panoramic roaming data, and the first virtual camera pose, and the second depth value of the second target vertex corresponding to the second pixel point in the second frame of panoramic roaming data, and the second virtual camera pose, a pose conversion relationship can be determined based on the first virtual camera pose and the second virtual camera pose; based on the determined pose conversion relationship, the first depth value is converted to the target first depth value corresponding to the second virtual camera pose, or the second depth value is converted to the target second depth value corresponding to the first virtual camera pose; after obtaining the target first depth value or the target second depth value, the first target vertex and the second target vertex can be determined based on the target first depth value and the second depth value. A first distance at depth; or, based on the first depth value and the target second depth value, determining the first distance between a target vertex and a second target vertex at depth; determining the distance between the pixel point corresponding to the first target vertex and the pixel point corresponding to the second target vertex when the first target vertex and the second target vertex are projected onto the same preset plane, and based on the distance and the camera projection principle, determining the second distance after the first target vertex and the second target vertex are projected onto the same preset plane; based on the first distance and the second distance, using the Pythagorean theorem, determining the third distance between the first target vertex and the second target vertex in the model coordinate system; based on the proportional relationship between the model coordinate system and the real space, and the third distance, determining the distance between the first target vertex and the second target vertex. The specific calculation method of determining the third distance between the first target vertex and the second target vertex in the model coordinate system using the Pythagorean theorem can be referred to Formula 1 in the specific implementation shown in C1 above, and the repeated parts will not be repeated.

[0126] In a specific implementation, after calculating the distance between the first target vertex and the second target vertex, the distance between the first target vertex and the second target vertex may be applied based on at least one of the following but not limited to D1 to D2:

[0127] D1. When the first target vertex belongs to the first object in the target space and the second target vertex belongs to the second object in the target space, the distance between the first object and the second object can be determined based on the distance between the first target vertex and the second target vertex.

[0128] The first object may include, for example, any target object in the target space; and the second object may include, for example, any target object in the target space except the first object.

[0129] Exemplarily, the distance l2 between the first target vertex and the second target vertex may be determined as the distance between the first object and the second object.

[0130] D2. When the first target vertex and the second target vertex belong to a first object in the target space, size information of the first object may be determined based on the distance between the first target vertex and the second target vertex.

[0131] Exemplarily, the target distance l2 between the first target vertex and the second target vertex may be determined as the size information of the first object.

[0132] Similarly, when the first target vertex and the second target vertex belong to a second object in the target space, the size information of the second object can be determined based on the distance between the first target vertex and the second target vertex. For example, the distance l2 between the first target vertex and the second target vertex can be determined as the size information of the second object.

[0133] In the embodiment of the present disclosure, a target three-dimensional model with pixel information of the target space is generated, and panoramic roaming data of the target space is generated by using the target three-dimensional model. The panoramic roaming data can be generated by directly performing pixel mapping based on the pixel information of each vertex in the target three-dimensional model, thereby reducing the distortion of objects in the target space in the generated panoramic roaming data and restoring the target space more realistically.

[0134] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0135] Based on the same inventive concept, a data processing device corresponding to the data processing method is also provided in the embodiment of the present disclosure. Since the principle of solving the problem by the device in the embodiment of the present disclosure is similar to the above-mentioned data processing method in the embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0136] Reference Figure 3 FIG. 1 is a schematic diagram of a data processing device provided by an embodiment of the present disclosure, wherein the device includes: an acquisition module 301, a first generation module 302, and a second generation module 303; wherein:

[0137] An acquisition module 301 is used to acquire image data of a target space; a first generation module 302 is used to generate a target three-dimensional model of the target space with pixel information based on the image data and the first posture of the image acquisition device in the target space when acquiring the image data; a second generation module 303 is used to generate panoramic roaming data based on the second posture of the virtual camera in the virtual scene corresponding to the target space and the target three-dimensional model in response to panoramic roaming being triggered.

[0138] In an optional embodiment, when the first generation module 302 generates a target three-dimensional model with pixel information in the target space based on the image data and the first pose of the image acquisition device in the target space when acquiring the image data, it is specifically used to: perform three-dimensional reconstruction of the target space based on the image data and the first pose of the image acquisition device in the target space when acquiring the image data to obtain a three-dimensional model of the target space; the three-dimensional model includes: multiple vertices on the surface of the target object located in the target space, and position information of each vertex in the target space; based on the pixel values ​​corresponding to each pixel point in the image data, pixel mapping is performed on the vertices corresponding to each pixel point to obtain pixel information corresponding to each vertex in the three-dimensional model; based on the pixel information corresponding to each vertex in the three-dimensional model, the target three-dimensional model is obtained.

[0139] In an optional implementation, the target object includes at least one of the following: a building in the target space, and equipment deployed in the target space.

[0140] In an optional embodiment, when the second generation module 303 generates panoramic roaming data in response to the panoramic roaming being triggered, based on the second posture of the virtual camera in the virtual scene corresponding to the target space and the target three-dimensional model, it is specifically used to: in response to the panoramic roaming being triggered, based on the second posture of the virtual camera in the virtual scene corresponding to the target space, determine the viewing angle range corresponding to the virtual camera in the second posture; based on the target three-dimensional model, generate a projection image in the viewing angle range; based on the projection image in the viewing angle range, generate the panoramic roaming data corresponding to the virtual camera in the second posture.

[0141] In an optional embodiment, when the second generation module 303 generates a projection image within the viewing angle range based on the target three-dimensional model, it is specifically used to: determine the projection plane corresponding to the virtual camera in the second posture; determine the target vertices from the target three-dimensional model based on the posture of the projection plane in the target three-dimensional model; project the pixel information corresponding to the target vertices onto the projection plane to generate a projection image within the viewing angle range.

[0142] In an optional embodiment, after executing the generation of the panoramic roaming data, the second generation module 303 is further used to: in response to the ranging operation of the first pixel point and the second pixel point in the panoramic roaming data, determine the distance between the first target vertex and the second target vertex based on the first depth value of the first target vertex corresponding to the first pixel point and the second depth value of the second target vertex corresponding to the second pixel point.

[0143] In an optional embodiment, the first pixel point and the second pixel point belong to the same frame of panoramic roaming data, and the second generation module 303 is further used to: determine the depth information corresponding to each target vertex based on the second posture of the virtual camera in the virtual scene corresponding to the target space; render the depth information corresponding to each target vertex into a preset cache; and in response to the ranging operation of the first pixel point and the second pixel point in the panoramic roaming data, read the first depth value corresponding to the first target vertex and the second depth value corresponding to the second target vertex from the preset cache.

[0144] In an optional embodiment, the first pixel point belongs to a first frame of panoramic roaming data, the second pixel point belongs to a second frame of panoramic roaming data, and the first frame of panoramic roaming data and the second frame of panoramic roaming data are different panoramic roaming data. The second generation module 303 is further used to: in response to a first pixel point selection operation on the first frame of panoramic roaming data, record a first virtual camera pose corresponding to the first frame of panoramic roaming data and a first depth value of the first target vertex corresponding to the first pixel point; in response to a second pixel point selection operation on the second frame of panoramic roaming data, record a second virtual camera pose corresponding to the second frame of panoramic roaming data and a second depth value of the second target vertex corresponding to the second pixel point.

[0145] In an optional embodiment, the second generation module 303, when executing the operation of selecting a first pixel point on the first frame of panoramic roaming data and recording the first virtual camera pose corresponding to the first frame of panoramic roaming data and the first depth value of the first target vertex corresponding to the first pixel point, is specifically used to: determine the depth information corresponding to each target vertex corresponding to the first frame of panoramic roaming data based on the first virtual camera pose of the virtual camera in the virtual scene corresponding to the target space when generating the first frame of panoramic roaming data; render the depth information corresponding to each target vertex into a preset cache; and read the first depth value of the first target vertex corresponding to the first pixel point from the preset cache in response to the operation of selecting a first pixel point on the first frame of panoramic roaming data.

[0146] In an optional embodiment, the second generation module 303, when executing the operation of selecting a second pixel point on the second frame of panoramic roaming data and recording the second virtual camera pose corresponding to the second frame of panoramic roaming data and the second depth value of the second target vertex corresponding to the second pixel point, is specifically used to: determine the depth information corresponding to each target vertex corresponding to the second frame of panoramic roaming data based on the second virtual camera pose of the virtual camera in the virtual scene corresponding to the target space when generating the second frame of panoramic roaming data; render the depth information corresponding to each target vertex into a preset cache; in response to the operation of selecting a second pixel point on the second frame of panoramic roaming data, read the second depth value of the second target vertex corresponding to the second pixel point from the preset cache, and record the second virtual camera pose and the second depth value of the second target vertex.

[0147] For descriptions of the processing flow of each module in the device and the interaction flow between each module, reference can be made to the relevant descriptions in the above method embodiment, which will not be described in detail here.

[0148] Based on the same technical concept, the embodiment of the present application also provides a computer device. Figure 4 4 is a schematic diagram of the structure of a computer device 400 provided in an embodiment of the present application, including a processor 401, a memory 402, and a bus 403. The memory 402 is used to store execution instructions and includes a memory 4021 and an external memory 4022. The memory 4021 is also referred to as internal memory and is used to temporarily store operation data in the processor 401 and data exchanged with an external memory 4022 such as a hard disk. The processor 401 exchanges data with the external memory 4022 through the memory 4021. When the computer device 400 is running, the processor 401 communicates with the memory 402 via the bus 403, so that the processor 401 executes the following instructions:

[0149] Acquire image data of a target space; generate a target three-dimensional model of the target space with pixel information based on the image data and a first pose of an image acquisition device in the target space when acquiring the image data; in response to panoramic roaming being triggered, generate panoramic roaming data based on a second pose of a virtual camera in a virtual scene corresponding to the target space and the target three-dimensional model.

[0150] The specific processing flow of the processor 401 can refer to the description of the above method embodiment and will not be repeated here.

[0151] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program executes the steps of the data processing method described in the above method embodiment. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0152] The embodiments of the present disclosure also provide a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the data processing method described in the above method embodiment. For details, please refer to the above method embodiment and will not be repeated here.

[0153] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0154] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0155] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0156] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0157] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a nonvolatile computer readable storage medium executable by a processor. Based on this understanding, the technical solutions of the present disclosure essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present disclosure. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0158] Finally, it should be noted that: the above-described embodiments are only specific embodiments of the present disclosure, used to illustrate the technical solutions of the present disclosure, and not to limit them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art who is familiar with the technology in the art can still make modifications or easily think of changes to the technical solutions described in the foregoing embodiments, or make equivalent replacements to some of the technical features; and these modifications, changes or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A data processing method, characterized in that: include: Acquire image data of the target space; generating a three-dimensional target model with pixel information in the target space based on the image data and a first pose of the image acquisition device in the target space when acquiring the image data; In response to the panoramic roaming being triggered, generating panoramic roaming data based on a second position of the virtual camera in the virtual scene corresponding to the target space and the target three-dimensional model; The step of generating a target three-dimensional model with pixel information in the target space based on the image data and a first pose of the image acquisition device in the target space when acquiring the image data comprises: Based on the image data and the first pose of the image acquisition device in the target space when acquiring the image data, the target space is three-dimensionally reconstructed to obtain a three-dimensional model of the target space; the three-dimensional model includes: a plurality of vertices on the surface of the target object located in the target space, and position information of each vertex in the target space; Based on the pixel values ​​corresponding to the respective pixel points in the image data, pixel mapping is performed on the vertices corresponding to the respective pixel points to obtain pixel information corresponding to the respective vertices in the three-dimensional model; Obtaining the target three-dimensional model based on pixel information corresponding to each vertex in the three-dimensional model; In response to the panoramic roaming being triggered, generating the panoramic roaming data based on the second position of the virtual camera in the virtual scene corresponding to the target space and the target three-dimensional model, includes: In response to the panoramic roaming being triggered, determining, based on a second posture of the virtual camera in the virtual scene corresponding to the target space, a viewing angle range corresponding to the second posture of the virtual camera; generating a projection image within the viewing angle range based on the target three-dimensional model; Based on the projection image within the viewing angle range, generating panoramic roaming data corresponding to the virtual camera in the second posture; Generating a projection image within the viewing angle range based on the target three-dimensional model includes: Determining a projection plane corresponding to the virtual camera in the second posture; Determining a target vertex from the target three-dimensional model based on the position of the projection plane in the target three-dimensional model; The pixel information corresponding to the target vertex is projected onto the projection plane to generate a projection image within the viewing angle range.

2. The method according to claim 1, characterized in that The target object includes at least one of the following: a building in a target space, and equipment deployed in the target space.

3. The method according to any one of claims 1-2, characterized in that After generating the panoramic roaming data, the method further includes: In response to a ranging operation on a first pixel point and a second pixel point in the panoramic roaming data, a distance between the first target vertex and the second target vertex is determined based on a first depth value of a first target vertex corresponding to the first pixel point and a second depth value of the second target vertex corresponding to the second pixel point.

4. The method according to claim 3, characterized in that The first pixel point and the second pixel point belong to the same frame of panoramic roaming data, and the method further includes: Determining depth information corresponding to each target vertex based on a second position of the virtual camera in the virtual scene corresponding to the target space; rendering the depth information corresponding to each target vertex into a preset cache; In response to the ranging operation on the first pixel point and the second pixel point in the panoramic roaming data, a first depth value corresponding to the first target vertex and a second depth value corresponding to the second target vertex are read from the preset cache.

5. The method according to claim 3, characterized in that The first pixel point belongs to a first frame of panoramic roaming data, the second pixel point belongs to a second frame of panoramic roaming data, and the first frame of panoramic roaming data and the second frame of panoramic roaming data are different panoramic roaming data. The method further includes: In response to performing a first pixel selection operation on the first frame of panoramic roaming data, recording a first virtual camera pose corresponding to the first frame of panoramic roaming data and a first depth value of a first target vertex corresponding to the first pixel; In response to a second pixel selection operation performed on the second frame of panoramic roaming data, a second virtual camera pose corresponding to the second frame of panoramic roaming data and a second depth value of a second target vertex corresponding to the second pixel are recorded.

6. The method according to claim 5, characterized in that The step of recording, in response to performing a first pixel selection operation on the first frame of panoramic roaming data, a first virtual camera pose corresponding to the first frame of panoramic roaming data and a first depth value of a first target vertex corresponding to the first pixel, includes: Determining, based on a first virtual camera pose of the virtual camera in the virtual scene corresponding to the target space when generating the first frame of panoramic roaming data, depth information corresponding to each target vertex corresponding to the first frame of panoramic roaming data; and rendering the depth information corresponding to each target vertex into a preset buffer; In response to a first pixel selection operation on the first frame panoramic roaming data, a first depth value of a first target vertex corresponding to the first pixel is read from the preset cache, and the first virtual camera pose and the first depth value of the first target vertex are recorded.

7. The method according to claim 5 or 6, characterized in that The step of recording a second virtual camera pose corresponding to the second frame of panoramic roaming data and a second depth value of a second target vertex corresponding to the second pixel point in response to performing a second pixel selection operation on the second frame of panoramic roaming data includes: Determining, based on a second virtual camera pose of the virtual camera in the virtual scene corresponding to the target space when generating the second frame of panoramic roaming data, depth information corresponding to each target vertex corresponding to the second frame of panoramic roaming data; and rendering the depth information corresponding to each target vertex into a preset buffer; In response to a second pixel selection operation on the second frame panoramic roaming data, a second depth value of a second target vertex corresponding to the second pixel is read from the preset cache, and the second virtual camera pose and the second depth value of the second target vertex are recorded.

8. A data processing device, characterized in that: include: An acquisition module is used to acquire image data of the target space; A first generating module is configured to generate a target three-dimensional model with pixel information in the target space based on the image data and a first pose of the image acquisition device in the target space when acquiring the image data; a second generating module, configured to generate panoramic roaming data based on a second position of the virtual camera in the virtual scene corresponding to the target space and the target three-dimensional model in response to the panoramic roaming being triggered; The step of generating a target three-dimensional model with pixel information in the target space based on the image data and a first pose of the image acquisition device in the target space when acquiring the image data comprises: Based on the image data and the first pose of the image acquisition device in the target space when acquiring the image data, the target space is three-dimensionally reconstructed to obtain a three-dimensional model of the target space; the three-dimensional model includes: a plurality of vertices on the surface of the target object located in the target space, and position information of each vertex in the target space; Based on the pixel values ​​corresponding to the respective pixel points in the image data, pixel mapping is performed on the vertices corresponding to the respective pixel points to obtain pixel information corresponding to the respective vertices in the three-dimensional model; Obtaining the target three-dimensional model based on pixel information corresponding to each vertex in the three-dimensional model; In response to the panoramic roaming being triggered, generating the panoramic roaming data based on the second position of the virtual camera in the virtual scene corresponding to the target space and the target three-dimensional model, includes: In response to the panoramic roaming being triggered, determining, based on a second posture of the virtual camera in the virtual scene corresponding to the target space, a viewing angle range corresponding to the second posture of the virtual camera; generating a projection image within the viewing angle range based on the target three-dimensional model; Based on the projection image within the viewing angle range, generating panoramic roaming data corresponding to the virtual camera in the second posture; Generating a projection image within the viewing angle range based on the target three-dimensional model includes: Determining a projection plane corresponding to the virtual camera in the second posture; Determining a target vertex from the target three-dimensional model based on the position of the projection plane in the target three-dimensional model; The pixel information corresponding to the target vertex is projected onto the projection plane to generate a projection image within the viewing angle range.

9. A computer device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is used to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the processor performs the steps of the data processing method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program. When the computer program is executed by a computer device, the computer device performs the steps of the data processing method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Three-dimensional space model reconstruction method and device and storage medium

    CN113570721A

  • Methods and apparatus for displaying and manipulating a panoramic image by tiles

    WO2014043814A1