Data processing method and device
By displaying multiple shooting modes on a handheld terminal, selecting the target shooting mode, and adapting the shooting strategy, the problem of resource waste caused by frequent scene updates in 3D space reconstruction is solved, achieving low-cost and efficient 3D space updates.
Patent Information
- Application Number
- CN202511215912.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-12-12
AI Technical Summary
In existing technologies for 3D spatial reconstruction, frequent scene updates require the re-capturing of a large number of images, resulting in high labor costs, waste of system and network resources, and a lack of low-cost and readily available shooting equipment.
By displaying multiple shooting modes on a handheld terminal, selecting the target shooting mode, and capturing two-dimensional images and depth information at the target location according to the shooting mode adaptation strategy, only necessary images are uploaded to update the three-dimensional virtual space model, and shooting is carried out in combination with low-cost handheld terminals and auxiliary equipment.
It reduces the labor costs, system resource and network resource consumption when frequently updating scenes, improves shooting convenience and flexibility, and reduces shooting costs.
Smart Images

Figure CN121120931A_ABST
Abstract
Description
[0001] This application is a divisional application of Chinese invention patent application filed on September 30, 2021, with application number 202111168457.X and title "A Data Processing Method and Apparatus". Technical Field
[0002] This application relates to the field of next-generation information technology, and in particular to a data processing method and apparatus. Background Technology
[0003] 3D reconstruction technology is a research hotspot in the field of computer vision, both in industry and academia. Depending on the object being reconstructed, it can be categorized into object 3D reconstruction, scene 3D reconstruction, and human body 3D reconstruction. In recent years, technologies for indoor scene 3D reconstruction have matured, giving rise to applications such as VR (Virtual Reality) house viewing and virtual shopping. These applications are primarily deployed on the Web, mainly using WebGL (Web Graphics Library) to render the 3D virtual space model. Summary of the Invention
[0004] This application discloses a data processing method and apparatus.
[0005] On one hand, this application discloses a data processing method applied to a handheld terminal, comprising: displaying at least one shooting mode on the screen of the handheld terminal; the shooting mode includes at least: a shooting mode of a handheld terminal with a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal without a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal with a distance sensor, a shooting mode of a handheld terminal without a distance sensor, and a shooting mode of a handheld terminal without a distance sensor combined with a panoramic camera; determining a selected target shooting mode among the displayed multiple shooting modes; and capturing a two-dimensional image of a three-dimensional space and depth information of the two-dimensional image according to a target shooting strategy adapted to the target shooting mode and a shooting device of the physical form corresponding to the target shooting mode.
[0006] In another aspect, this application discloses a data processing apparatus applied to a handheld terminal, comprising: a display module for displaying at least one shooting mode on the screen of the handheld terminal; the shooting modes include at least: a shooting mode of a handheld terminal with a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal without a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal with a distance sensor, a shooting mode of a handheld terminal without a distance sensor, and a shooting mode of a handheld terminal without a distance sensor combined with a panoramic camera; a second determining module for determining a selected target shooting mode among the displayed multiple shooting modes; and a shooting module for capturing a two-dimensional image of three-dimensional space and depth information of the two-dimensional image according to a target shooting strategy adapted to the target shooting mode and a shooting device of the physical form corresponding to the target shooting mode.
[0007] In another aspect, this application discloses an electronic device comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to perform a data processing method as described in any of the preceding aspects.
[0008] In another aspect, this application discloses a non-transitory computer-readable storage medium that, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to perform the data processing method as described in any of the preceding aspects.
[0009] In another aspect, this application discloses a computer program product in which, when the instructions in the computer program product are executed by a processor of an electronic device, the electronic device is enabled to perform the data processing method as described in any of the preceding aspects.
[0010] Compared with the prior art, this application has the following advantages: In this application, at least one shooting mode is displayed on the screen of a handheld terminal; the shooting modes include at least: a shooting mode of a handheld terminal with a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal without a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal with a distance sensor, a shooting mode of a handheld terminal without a distance sensor, and a shooting mode of a handheld terminal without a distance sensor combined with a panoramic camera; determining the selected target shooting mode among the displayed multiple shooting modes; and capturing a two-dimensional image of three-dimensional space and depth information of the two-dimensional image according to a target shooting strategy adapted to the target shooting mode and a shooting device with the physical form corresponding to the target shooting mode.
[0011] This application enables the use of a handheld terminal to capture two-dimensional images in three-dimensional space when such images are required, eliminating the need for specialized shooting equipment. Due to the low cost and high ownership rate of handheld terminals, shooting costs can be reduced and shooting convenience can be improved. Attached Figure Description
[0012] Figure 1 This is a schematic diagram of the structure of a data processing system according to this application.
[0013] Figure 2 This is a flowchart of the steps of a data processing method according to this application.
[0014] Figure 3 This is a flowchart of the steps of a data processing method according to this application.
[0015] Figure 4 This is a flowchart of the steps of a data processing method according to this application.
[0016] Figure 5 This is a schematic diagram of the structure of a depth information acquisition model according to this application.
[0017] Figure 6 This is a schematic diagram of the structure of a fusion layer in this application.
[0018] Figure 7 This is a schematic diagram of the structure of a fusion layer in this application.
[0019] Figure 8 This is a schematic diagram of the structure of a fusion layer in this application.
[0020] Figure 9 This is a structural block diagram of a data processing device according to this application.
[0021] Figure 10 This is a structural block diagram of a data processing device according to this application.
[0022] Figure 11 This is a structural block diagram of a data processing device according to this application.
[0023] Figure 12 This is a structural block diagram of a data processing device according to this application.
[0024] Figure 13 This is a structural block diagram of a device according to this application. Detailed Implementation
[0025] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] Reference Figure 1 The diagram illustrates the structure of a data processing system according to this application. The system includes a cloud platform 01 and at least one camera terminal 02. Each camera terminal 02 is communicatively connected to the cloud platform 01.
[0027] The shooting terminal can be a terminal used by the shooting party (such as the organizer or exhibitor of an exhibition). For example, a shooting party might be setting up a 3D space for an exhibition. Within this space, the shooting party can arrange its own venue and place at least one of its own items (such as exhibits) so that visitors (customers) can view these items. Furthermore, if a user is interested in a particular item, they can purchase it from the shooting party.
[0028] In one example, the photographer may include the seller, etc. The photographer's terminal may include the photographer's handheld terminal, such as a mobile phone, tablet, or camera, which at least has image acquisition capabilities.
[0029] In one approach, when it is necessary to generate a three-dimensional virtual space model, the camera can use its terminal to capture two-dimensional images of the three-dimensional space from multiple shooting positions in the three-dimensional space (the two-dimensional images also include depth information, etc.). The camera terminal can acquire the two-dimensional images of the three-dimensional space captured from multiple shooting positions in the three-dimensional space, and generate topological relationships, which include: multiple shooting positions and the interdependencies between the multiple shooting positions.
[0030] The interdependence between multiple shooting locations can be understood as follows: any one shooting location is interdependent with at least one other shooting location. For example, if one shooting location depends on another, it means that in the process of generating a 3D virtual space model, a 2D image taken at one shooting location can be stitched together with a 2D image taken at the other shooting location.
[0031] Topological relationships can be represented as a directed graph. For example, if one shooting position depends on another, the topological relationship is represented by arrows pointing from the other shooting position to the first shooting position. The camera operator can then send 2D images of 3D space taken from multiple shooting positions in 3D space, along with the topological relationships, to the cloud. The cloud receives these images and generates a 3D virtual space model based on the data, and stores at least this model. Subsequently, when a user terminal requests the 3D virtual space model, the cloud can send it to the user terminal for display. The user can then virtually navigate and view 3D space within the 3D virtual space model.
[0032] This allows users to view the 3D space even when they cannot go to the actual 3D space, freeing them from the time and location restrictions of the 3D space. This enables users who have not gone to the 3D space to view it, and also allows them to view the 3D space outside of its continuous existence time, thereby increasing the promotion rate of items placed in the 3D space.
[0033] However, in one possible scenario, the two-dimensional images and topological relationships of the three-dimensional space captured at multiple shooting locations in the three-dimensional space are obtained statically in a single instance. After the three-dimensional virtual space model of the three-dimensional space is generated in the cloud, neither the shooting terminal nor the cloud usually retains the two-dimensional images and topological relationships of the three-dimensional space captured at multiple shooting locations in the three-dimensional space.
[0034] Subsequently, if changes occur in the scene arrangement within the 3D space, requiring the 3D virtual space model to adapt to the changed parts of the 3D space, one approach involves the user using a camera terminal to re-capture 2D images of the 3D space from multiple shooting positions. The camera terminal acquires these 2D images from the multiple shooting positions, generates a topological relationship, and then sends the images and topological relationship to the cloud. The cloud receives the images and topological relationship from the multiple shooting positions and regenerates the 3D virtual space model based on them.
[0035] However, the inventors discovered the following problems with the above method: First, the user needs to use the shooting terminal to take two-dimensional images of the three-dimensional space from multiple shooting positions in the three-dimensional space, resulting in high labor costs. Second, the shooting terminal needs to acquire two-dimensional images of the three-dimensional space from multiple shooting positions in the three-dimensional space again, which wastes the terminal's system resources. Third, sending the two-dimensional images of the three-dimensional space taken from multiple shooting positions in the three-dimensional space and the topology relationship to the cloud again wastes network resources.
[0036] Therefore, in order to solve the above-mentioned technical problems, refer to Figure 2 This document illustrates a flowchart of a data processing method according to this application. In this method, if a 3D virtual space model adapts to the changed portion of the 3D scene due to changes in the scene arrangement within the 3D space, the photographer can capture a new 2D image of the 3D space at the target shooting location where the shooting content needs to be changed. Based on the original topological relationship, the photographer obtains the target topological relationship related to the target shooting location and uploads the new 2D image and the target topological relationship to the cloud. This allows the cloud to generate a new 3D virtual space model based on the target topological relationship and the new 2D image. This method is applied to… Figure 1 In the cloud 01 and one of the shooting terminals 02 shown, the method may specifically include the following steps: In step S101, the shooting terminal obtains the original topology relationship, which includes: multiple shooting positions in the three-dimensional space involved in the previous establishment of the three-dimensional virtual space model and the interdependence between the multiple shooting positions.
[0037] In this application, the original topology is generated by the camera terminal during the previous establishment of the three-dimensional virtual space model and cached in the camera terminal or in the cloud. Therefore, the camera terminal can obtain the original topology cached by the camera terminal or the original topology cached in the cloud.
[0038] In this application, if the 3D virtual space model needs to be adapted to the changed part of the 3D space due to changes in some scenes arranged in the 3D space, the shooting party can input a change request on the shooting party's terminal to change the 3D virtual space model of the 3D space by "taking a new 2D image of the 3D space at the target shooting position where the shooting content needs to be changed". The shooting party receives the change request and can then execute step S101.
[0039] In step S102, the shooting terminal determines the target shooting location from among multiple shooting locations in the original topology relationship, which is the content to be changed.
[0040] The original topology includes multiple shooting locations in three-dimensional space. Sometimes, the photographer needs to directly input a specified operation on their terminal to specify the target shooting location from among the multiple shooting locations in the original topology for the content to be changed. Based on this specified operation, the photographer's terminal can determine the target shooting location from among the multiple shooting locations in the original topology for the content to be changed.
[0041] The target shooting location can be an existing location among multiple shooting locations. Alternatively, at least some of the target shooting locations can be locations that do not exist among multiple shooting locations, that is, new shooting locations.
[0042] In step S103, the shooting terminal obtains the target interdependence relationship between the target shooting position and the shooting position in the original topology relationship based on the original topology relationship and the target shooting position.
[0043] Target dependencies can include only the dependencies between the target location and nearby shooting locations in the original topology relationship, and may exclude dependencies between shooting locations that are not associated with the target location (far away). The target shooting location and its dependencies can be represented by target topology relationships related to the target shooting location, which are obtained based on the original topology relationship and the target configuration location. The target topology relationship contains the dependencies between the target shooting location and the shooting locations in the original topology relationship.
[0044] In one embodiment of this application, if the target shooting location exists in the original topological relationship, the original topological relationship can be determined as the target topological relationship.
[0045] The existence of a target shooting location in the original topology can be understood as follows: when there is only one target shooting location, this target shooting location is the same as one of the multiple shooting locations in the original topology. When there are at least two target shooting locations, each target shooting location is the same as one of the multiple shooting locations in the original topology.
[0046] Alternatively, if the target shooting location does not exist in the original topology, the original topology is updated based on at least the target shooting location, the associated shooting locations in the original topology that are interdependent with the target shooting location, and the dependency between the target shooting location and the associated shooting locations, to obtain the target topology.
[0047] The statement that a target shooting location does not exist in the original topology can be understood as follows: when there is only one target shooting location, this single target shooting location is different from every single shooting location in the original topology. When there are at least two target shooting locations, at least one target shooting location is different from every single shooting location in the original topology.
[0048] In the case of a single target shooting location, the original topology can be updated based on the target shooting location, the associated shooting locations in the original topology that are mutually dependent on the target shooting location, and the dependency between the target shooting location and the associated shooting locations to obtain the target topology.
[0049] Alternatively, if there are at least two target shooting locations, there may be dependencies between the at least two target shooting locations. Therefore, the original topology can be updated based on the target shooting locations, the associated shooting locations that are mutually dependent on the target shooting locations in the original topology, the dependencies between the target shooting locations and the associated shooting locations, and the dependencies between the at least two target shooting locations to obtain the target topology.
[0050] Among them, the dependency relationship between the target shooting location and the associated shooting location can be manually set by the shooting party, and the dependency relationship between at least two target shooting locations can also be manually set by the shooting party, etc.
[0051] In step S104, the shooting terminal acquires a new two-dimensional image of the three-dimensional space captured at the target shooting location.
[0052] In this application, when there is only one target shooting location, the shooting party can control the shooting terminal to take a picture at that target shooting location, and the shooting terminal can obtain a new two-dimensional image of the three-dimensional space captured at that target shooting location. When there are at least two target shooting locations, the shooting party can control the shooting terminal to take pictures sequentially at each target shooting location, and the shooting terminal can sequentially obtain new two-dimensional images of the three-dimensional space captured at each target shooting location.
[0053] In this step, there is no restriction on the execution order between steps S103 and S104. Steps S103 and S104 can be executed concurrently, or steps S103 can be executed first and then steps S104 can be executed, or steps S104 can be executed first and then steps S103 can be executed.
[0054] In step S105, the shooting terminal sends the target shooting location, the interdependence of the targets, and a new two-dimensional image to the cloud.
[0055] In step S106, the cloud receives the target shooting location, the interdependence of the targets, and the new two-dimensional image sent by the shooting terminal.
[0056] In step S107, the cloud retrieves cached old two-dimensional images, which include: two-dimensional images taken at shooting locations other than the target shooting location, which were involved in the previous establishment of the three-dimensional virtual space model.
[0057] In this application, the old two-dimensional images were acquired by the shooting terminal during the previous process of establishing a three-dimensional virtual space model and sent to the cloud, and are cached in the cloud. Therefore, the shooting terminal can obtain the old two-dimensional images cached in the cloud.
[0058] In step S108, the cloud generates a new three-dimensional virtual space model based on the target shooting location, the interdependence of the targets, the new two-dimensional image, and the old two-dimensional image.
[0059] In this application, the three-dimensional virtual space model includes VR (Virtual Reality) models, MR (Mixed Reality) models, and AR (Augmented Reality) models, etc. Of course, it may also include other forms of models. This application does not limit the specific model.
[0060] The cloud can store new three-dimensional virtual space models in the cloud, or store new three-dimensional virtual space models in other storage devices associated with the cloud, so that user terminals can call the new three-dimensional virtual space models later.
[0061] For example, when a user terminal requests a new 3D virtual space model, the cloud can send the new 3D virtual space model to the user terminal for display. The user can then virtually navigate within the new 3D virtual space model, viewing the 3D space and objects displayed within it. In this application, "objects" include salable goods.
[0062] In one embodiment of this application, the cloud can store the identification information of the shooting party and the new three-dimensional virtual space model generated in the correspondence between identification information and three-dimensional virtual space model.
[0063] Furthermore, new 2D images can be cached in the cloud. For example, new 2D images can be cached in the cloud so that if, later, due to changes in the arrangement of parts of the scene in 3D space, the 3D virtual space model needs to be adapted to the changed parts of the scene in 3D space, the new 2D images can be directly obtained, and at least the 3D virtual space model can be regenerated using the above method with reference to the new 2D images.
[0064] In this application, the shooting terminal obtains the original topology, which includes multiple shooting positions in the three-dimensional space involved in the previous establishment of the three-dimensional virtual space model and the interdependencies between these shooting positions. The target shooting position, containing the content to be changed, is determined from among the multiple shooting positions in the original topology. Based on the original topology and the target shooting position, the target interdependencies between the target shooting position and the shooting positions in the original topology are obtained; and a new two-dimensional image of the three-dimensional space captured at the target shooting position is acquired; the target shooting position, the target interdependencies, and the new two-dimensional image are then sent to the cloud.
[0065] The cloud receives the target shooting location, target dependencies, and new 2D images sent by the shooting terminal. It also retrieves cached older 2D images, including those taken at multiple shooting locations other than the target shooting location, which were involved in the previous creation of the 3D virtual space model. Based on the target shooting location, target dependencies, new 2D images, and older 2D images, a new 3D virtual space model is generated.
[0066] Through this application, if the 3D virtual space model needs to adapt to the changed parts of the 3D space due to changes in some scenes arranged in the 3D space, the shooting party does not need to use the 2D images of the 3D space taken by the shooting party terminal at each required shooting position in the 3D space again, the shooting party terminal does not need to acquire the 2D images of the 3D space taken at each required shooting position in the 3D space again, and does not need to send the 2D images of the 3D space taken at each required shooting position in the 3D space to the cloud again. The camera operator can use its terminal to capture new 2D images of the target location in 3D space, focusing solely on the content to be captured. The terminal can then send only these new 2D images to the cloud. The cloud can then use these new images, along with previously existing ones, to generate a new 3D virtual space model. This reduces the amount of data required for updates, thus lowering labor costs and conserving system and network resources for the camera operator. The more frequently the 3D space is updated, the more significant the savings in labor, system, and network resources become.
[0067] In another embodiment of this application, when the shooting terminal obtains the target topology relationship, the target topology relationship can also be cached. For example, the target topology relationship can be cached locally on the shooting terminal, and / or cached in the cloud, so that when "if the 3D virtual space model needs to adapt to the changed part of the 3D space due to changes in some scenes arranged in the 3D space", the cached target topology relationship can be directly obtained, and the target topology relationship can be used as the original topology relationship, and the 3D virtual space model of the 3D space can be regenerated with reference to the above method.
[0068] Furthermore, based on the interdependence between the various target shooting positions, it can output guidance information regarding the shooting order between the various target shooting positions, so that the photographer can take new two-dimensional images of three-dimensional space in sequence at each target shooting position according to the guidance information, which facilitates the photographer's shooting.
[0069] In another embodiment of this application, when there is only one target shooting location, the camera can control the camera terminal to take a picture at that single target shooting location to obtain a new two-dimensional image in three-dimensional space. When there are at least two target shooting locations, the camera can control the camera terminal to take pictures sequentially at each target shooting location to obtain new two-dimensional images in three-dimensional space sequentially.
[0070] When a new two-dimensional image of three-dimensional space is captured at a target shooting location, the new two-dimensional image can be displayed on the screen of the shooting terminal in real time, so that the shooting party can preview the new two-dimensional image captured at the target shooting location and check whether the new two-dimensional image captured matches the actual scene that needs to be captured at the target shooting location.
[0071] If the newly captured 2D image matches the scene actually needed to be captured at the target shooting location, the photographer can control the camera to continue shooting at the next target shooting location. If the newly captured 2D image does not match the scene actually needed to be captured at the target shooting location, the photographer needs to control the camera to reshoot at that target shooting location (sometimes this may require fine-tuning the target shooting location, for example, the photographer making minor adjustments to the target shooting location), until the newly captured 2D image matches the scene actually needed to be captured at the target shooting location.
[0072] Specifically, when the photographer needs to control their terminal to retake a shot at a target shooting location, the photographer can input a retake command into their terminal based on a newly displayed 2D image. The shooting terminal receives the retake command input based on the newly displayed 2D image. Based on the shooting command, a new 2D image in three-dimensional space is retaken from the target shooting location.
[0073] Here, "reshooting at a target location" can be understood as: reshooting at the target location, or, after making a minor adjustment to the target location, shooting at the adjusted target location, etc.
[0074] In this application, sometimes there are two or more target shooting positions, and at least one target shooting position does not exist in the original topology. That is, at least one target shooting position is a new shooting position in the three-dimensional space selected by the shooting party. After the new shooting position is set in the original topology, there is no dependency relationship between the new shooting position and the shooting position that already exists in the original topology. It is often necessary for the shooting party to set the dependency relationship between the new shooting position and the shooting position that already exists in the original topology in order to obtain the target topology.
[0075] However, sometimes the filming crew may overlook certain dependencies when setting up new shooting positions in the original topology, resulting in some new shooting positions being isolated. This means that the new shooting positions have no dependencies on other shooting positions, which will affect the accuracy of the new 3D virtual space model generated in the cloud. For example, it may affect the accuracy of the scene (or image) related to the new shooting position in the new 3D virtual space model generated in the cloud.
[0076] Therefore, in order to solve the above-mentioned technical problems, in another embodiment of this application, when the target topology is obtained, the shooting terminal can detect whether there are isolated shooting positions in the target topology. For example, it can detect whether there are shooting positions in the target topology that are not dependent on other shooting positions. If there are shooting positions that are not dependent on other shooting positions, they can be regarded as isolated shooting positions. If there are no shooting positions that are not dependent on other shooting positions, it can be determined that there are no isolated shooting positions in the target topology.
[0077] In cases where isolated shooting locations exist, a prompt message is output, such as displayed on the shooting terminal. This prompt message indicates the isolated shooting location within the target topology. Based on the prompt message, the shooting party can identify the isolated shooting location within the target topology and then input dependency setting operations according to the prompt message. These setting operations are used to establish dependencies between the isolated shooting location and other shooting locations besides the isolated one.
[0078] For the shooting terminal, it can obtain the dependency setting operation based on the prompt information, and set the dependency relationship between isolated shooting positions in the target topology and other shooting positions in the target topology. This eliminates isolated shooting positions in the target topology.
[0079] In one scenario, the data required to render a 3D virtual space model includes: a textured 3D model of an indoor scene, 360° panoramic views of the indoor scene taken from multiple shooting positions, and then a 3D virtual space model of the indoor scene is generated based on the 360° panoramic views from multiple shooting positions and the textured 3D model of the indoor scene.
[0080] In order to obtain a 360° panoramic view of any shooting location, it is necessary to capture multiple two-dimensional images from different perspectives and the depth information of each two-dimensional image at that shooting location. In order to obtain the depth information of the two-dimensional images, one method is to use professional equipment to capture the two-dimensional images and obtain the depth information of the two-dimensional images, such as depth cameras or LiDAR.
[0081] However, professional equipment such as depth cameras or LiDAR is often difficult to obtain. Professional equipment often needs to be purchased or rented separately, but the selling or renting price of professional equipment is often high, resulting in high shooting costs and low shooting convenience.
[0082] Therefore, in order to solve the above-mentioned technical problems, refer to Figure 3 The diagram illustrates a flowchart of a data processing method according to this application, which supports shooting using a low-cost and readily available handheld terminal. This method is applied to... Figure 1 In one of the shooting terminals 02 shown, the method may specifically include the following steps: In step S201, at least one shooting mode is displayed on the screen of the handheld terminal; the shooting modes include at least: a shooting mode of a handheld terminal with a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal without a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal with a distance sensor, a shooting mode of a handheld terminal without a distance sensor, and a shooting mode of a handheld terminal without a distance sensor combined with a panoramic camera.
[0083] In this application, the handheld terminals on the market include: handheld terminals with distance sensors and handheld terminals without distance sensors. When shooting, the handheld terminal held by the photographer may be a handheld terminal with a distance sensor or a handheld terminal without a distance sensor, and the photographer may or may not hold a stabilizer, and may or may not hold a panoramic camera.
[0084] Therefore, in order to support the shooting party in capturing two-dimensional images of three-dimensional space and corresponding depth information in various situations, the handheld terminal is equipped with an application for shooting. The application supports multiple shooting modes, including at least: shooting mode of a handheld terminal with a distance sensor mounted on a stabilizer, shooting mode of a handheld terminal without a distance sensor mounted on a stabilizer, shooting mode of a handheld terminal with a distance sensor, shooting mode of a handheld terminal without a distance sensor, and shooting mode of a handheld terminal without a distance sensor combined with a panoramic camera.
[0085] To improve the quality of the captured 2D images and corresponding depth information, thereby enhancing the quality of the generated 3D virtual space model, the application incorporates shooting strategies adapted to different shooting modes. Details of these shooting strategies will be provided later and will not be elaborated upon here.
[0086] In one scenario, the photographer has their own physical shooting device and can select the corresponding shooting mode on the application. For example, the photographer can control a handheld terminal to launch the application and control the application to display a shooting mode selection page. The shooting mode selection page displays at least one shooting mode, and the photographer can choose a suitable shooting mode from multiple options.
[0087] In step S202, the selected target shooting mode is determined from at least one of the displayed shooting modes.
[0088] In step S203, a two-dimensional image of the three-dimensional space and the depth information of the two-dimensional image are captured by the target shooting strategy adapted to the target shooting mode and the shooting device corresponding to the physical form of the target shooting mode.
[0089] At least one shooting mode is displayed on the screen of a handheld terminal; the shooting modes include at least: a shooting mode of a handheld terminal with a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal without a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal with a distance sensor, a shooting mode of a handheld terminal without a distance sensor, and a shooting mode of a handheld terminal without a distance sensor combined with a panoramic camera; a selected target shooting mode is determined from the at least one displayed shooting modes; a two-dimensional image of three-dimensional space and depth information of the two-dimensional image are captured according to the target shooting strategy adapted to the target shooting mode and the shooting device corresponding to the physical form of the target shooting mode.
[0090] The shooting devices corresponding to the shooting modes include: handheld terminals, or a combination of handheld terminals and panoramic cameras, or a stabilizer. The handheld terminals support handheld terminals without distance sensors or handheld terminals with distance sensors.
[0091] This application enables the use of a handheld terminal to capture two-dimensional images in three-dimensional space when such images are required, eliminating the need for specialized shooting equipment. Due to the low cost and high ownership rate of handheld terminals, this application reduces shooting costs and improves shooting convenience.
[0092] When auxiliary accessories are needed to improve shooting quality, accessories such as stabilizers and panoramic cameras are readily available and inexpensive. Compared to using expensive professional equipment, this application can save costs while requiring the same or similar shooting quality.
[0093] Secondly, when the photographer has a handheld terminal, which may or may not have a distance sensor, and the photographer may or may not have a stabilizer, and may or may not have a panoramic camera, this application can support multiple shooting modes so that the photographer can choose the appropriate shooting mode when needed. For example, the photographer can choose the appropriate shooting mode based on the physical shooting equipment they currently have, thereby improving the flexibility of shooting.
[0094] In one embodiment of this application, for the shooting mode of a handheld terminal with a distance sensor mounted on a stabilizer, the shooting strategy adapted to the shooting mode and the shooting device corresponding to the shooting mode capture a two-dimensional image of three-dimensional space and the depth information of the two-dimensional image, as detailed below: Distance sensors include LiDAR (Light Detection and Ranging) sensors. LiDAR sensors offer high accuracy, thus improving the accuracy of subsequently created panoramic images of 3D space and the accuracy of 3D virtual space models. Of course, other distance sensors may also be included; this application does not limit the specific model or type of sensor.
[0095] The handheld terminal can have an application installed, which may include an application corresponding to the stabilizer. The application can act as a communication interface between the handheld terminal and the stabilizer. That is, by operating the application on the handheld terminal, data interaction can be achieved between the handheld terminal and the stabilizer, such as controlling the operation of the stabilizer. Controlling the operation of the stabilizer includes controlling the rotation of the stabilizer (including rotation along the horizontal direction and rotation along the vertical direction). The handheld terminal and the stabilizer can communicate and connect based on Bluetooth, WiFi (Wireless Fidelity), or NFC (Near Field Communication).
[0096] The handheld device has a shooting function; for example, it has a camera, which enables it to shoot. The stabilizer includes two-axis, three-axis, four-axis, or five-axis stabilizers, etc. Furthermore, the shooter can hold the stabilizer by hand or place it on a tripod, etc.
[0097] Thus, in this application, the photographer can manually select multiple shooting positions in three-dimensional space. The distance between each shooting position can be 5 meters, 6 meters, or 7 meters, or other values. This application does not limit the distance between each shooting position.
[0098] The photographer can place a handheld device equipped with a distance sensor mounted on a stabilizer at one of the shooting positions. Then, the user controls the stabilizer's rotation (both horizontal and vertical) via an application on the handheld device. During the stabilizer's rotation, the handheld device's camera continuously captures images, including at least two frames, each depicting a scene in three-dimensional space. Thus, each of the at least two frames can be considered a two-dimensional image within the three-dimensional space. Furthermore, while capturing two-dimensional images, the handheld device can also obtain depth information from the captured two-dimensional images using the distance sensor. Therefore, at this shooting position, at least two two-dimensional images and the depth information for each image can be obtained.
[0099] However, noise can sometimes affect the accuracy of depth information in a single 2D image. Therefore, to improve the accuracy of depth information, in another embodiment, TSDF (Truncated Signed Distance Function, a surface reconstruction algorithm that uses structured point cloud data and parametrically represents surfaces) technology can be used to fuse the depth information of each 2D image obtained at the shooting location, resulting in fused depth information. Then, based on the internal parameters of the handheld terminal and the differences in pose information of the handheld terminal when each 2D image was captured at each shooting location, the fused depth information can be rendered to obtain the rendered depth information of each 2D image.
[0100] The fused depth information can be a depth map, which can be represented by a 3D white model, etc. The depth map includes multiple triangular faces. Based on the internal parameters of the handheld terminal and the differences in the pose information of the handheld terminal when capturing various 2D images at different shooting positions, each triangular face on the depth map is projected onto the image plane, and triangular faces that exceed the image range are removed.
[0101] Triangulation is performed on each projection point on the image plane to obtain multiple triangular patches on the image plane.
[0102] For any triangular patch in the depth map, assuming the triangular patch is △ABC, the depth values dA, dB, and dC of 3D points A, B, and C can be obtained by projecting them onto the image plane.
[0103] For any pixel P within triangle ABC, the depth value of pixel P can be calculated as follows: .
[0104] Where S△PBC is the area of △PBC, S△PAC is the area of △PAC, S△PAB is the area of △PAB, and S△ABC is the area of △ABC.
[0105] The same applies to every other pixel within triangle ABC.
[0106] The same operation is performed for each other triangular facet in the depth map.
[0107] In this process, the rendered depth information minimizes the impact of noise, thereby improving the accuracy of the depth information in each of the resulting two-dimensional images.
[0108] In addition, the internal parameters of the handheld terminal include the focal length of the handheld terminal, which can remain unchanged during the process of capturing a two-dimensional image at this shooting position.
[0109] In one example, the homography matrix between two adjacent two-dimensional images captured at the shooting location can be obtained, and then the focal length of the handheld terminal can be decomposed from the homography matrix. For specific decomposition methods, please refer to existing methods, which will not be detailed here.
[0110] Secondly, the pose information of the handheld terminal when capturing each two-dimensional image at the shooting position can be determined based on the rotation angle of the stabilizer when capturing each two-dimensional image. For example, assuming that three two-dimensional images are captured sequentially at the shooting position, namely two-dimensional image A, two-dimensional image B, and two-dimensional image C, the rotation angle of the stabilizer is 0° when capturing two-dimensional image A, 45° when capturing two-dimensional image B, and 90° when capturing two-dimensional image C.
[0111] The difference between a rotation angle of 45° and a rotation angle of 0° is 45°, and the difference between a rotation angle of 90° and a rotation angle of 45° is also 45°. Therefore, 45° can be used as the difference information of the pose of the handheld terminal when capturing various two-dimensional images at this shooting position.
[0112] After capturing a 2D image at the shooting location, the handheld terminal with a distance sensor mounted on the stabilizer can be placed at other shooting locations, and the above process can be performed at other shooting locations to obtain at least two 2D images and depth information of each 2D image at other shooting locations, until 2D images are captured and depth information of the 2D images is obtained at each shooting location.
[0113] By combining the handheld terminal and the stabilizer in the above manner for joint control, the stabilizer's rotational movement is used to capture 360° x 180° images (all-around image capture, including the entire spherical surface). Due to the participation of the stabilizer, the influence of human shaking on the captured image can be avoided during the image capture process, and the center of the image capturing terminal can be located on the rotation axis of the stabilizer, so that the 360° x 180° panoramic image with the lowest possible level of defects or even no defects can be generated.
[0114] In addition, the above method only requires a stabilizer and a handheld terminal with a distance sensor, so the hardware cost of the above method is low.
[0115] Furthermore, in another embodiment of this application, for at least two two-dimensional images captured at any shooting location, the at least two two-dimensional images can be stitched together using a back-projection method (overlapping areas can be merged during the stitching process) to obtain a panoramic image of the three-dimensional space based on the shooting location. For example, if the at least two two-dimensional images are captured during a pure rotational movement of a handheld terminal at the shooting location, each pixel in the at least two two-dimensional images can be projected onto the corresponding latitude and longitude map according to the viewing direction to obtain a panoramic image of the three-dimensional space based on the shooting location.
[0116] Alternatively, at least two venue images can be stitched together using optical flow (overlapping areas can be merged during the stitching process) to obtain a panoramic view of the exhibition venue based on the acquisition location (this panoramic view obtained using optical flow has a low degree of imperfection).
[0117] The same process is repeated for at least two 2D images taken at each of the other shooting locations, resulting in a panoramic view of the 3D space based on each shooting location. This panoramic view of the 3D space from each shooting location can then be displayed for exhibitors to preview.
[0118] The back projection method yields panoramic images at a high rate, allowing the photographer to preview the panoramic image in a timely manner and determine whether the content in the panoramic image meets the requirements. If not, the image can be reshot promptly.
[0119] In one embodiment of this application, for the shooting mode of a handheld terminal with a distance sensor, the shooting strategy adapted to the shooting mode and the shooting device corresponding to the shooting mode capture a two-dimensional image of three-dimensional space and the depth information of the two-dimensional image, as detailed below: Distance sensors include LiDAR sensors, which offer high accuracy, thus improving the accuracy of subsequently created panoramic images of 3D space and the accuracy of 3D virtual space models. Other distance sensors may also be included; this application does not limit the specific model or type of sensor.
[0120] The handheld terminal has a shooting function, for example, the handheld terminal has a camera, which enables the handheld terminal to take pictures.
[0121] Thus, in this application, the photographer can manually select multiple shooting positions in three-dimensional space. The distance between each shooting position can be 5 meters, 6 meters, or 7 meters, or other values. This application does not limit the distance between each shooting position.
[0122] The handheld terminal can have an application installed for taking pictures. The person taking the picture can hold the handheld terminal with a distance sensor and stand at one of the shooting positions to take pictures.
[0123] During the shooting process, the handheld terminal can generate shooting guidance information (such as a circle) and a shooting crosshair (such as a ring) on the screen based on the data from the gyroscope in the handheld terminal. The user can move the position of the handheld terminal in space and change the orientation of the handheld terminal's camera to adjust the shooting crosshair so that the shooting crosshair coincides with the shooting guidance information. Then, the user can control the camera to capture a frame of image and then rotate the handheld terminal horizontally.
[0124] During the rotation of the handheld terminal, the handheld terminal can generate another shooting guidance information and shooting crosshair on the screen based on the data from the gyroscope in the handheld terminal. The user can move the position of the handheld terminal in space and change the orientation of the handheld terminal's camera to adjust the shooting crosshair so that the shooting crosshair coincides with the shooting guidance information. Then, the user can control the camera to take another frame of image, and so on. The captured images include at least two frames of images, and each of the at least two frames of images includes a three-dimensional space scene. Thus, each of the at least two frames of images can be regarded as a two-dimensional image in three-dimensional space.
[0125] Secondly, when a handheld terminal captures a 2D image, it can also acquire depth information of the captured 2D image through a distance sensor. Thus, at that shooting location, at least two 2D images and depth information for each image can be obtained.
[0126] However, noise can sometimes affect the accuracy of depth information in a single 2D image. Therefore, to improve the accuracy of depth information, in another embodiment, TSDF technology can be used to fuse the depth information of various 2D images obtained at the shooting location to obtain fused depth information. Then, based on the internal parameters of the handheld terminal and the differences in pose information of the handheld terminal when each 2D image was captured at each shooting location, the fused depth information can be rendered to obtain the rendered depth information of each 2D image.
[0127] The fused depth information can be a depth map, which can be represented by a 3D white model, etc. The depth map includes multiple triangular faces. Based on the internal parameters of the handheld terminal and the differences in the pose information of the handheld terminal when capturing various 2D images at different shooting positions, each triangular face on the depth map is projected onto the image plane, and triangular faces that exceed the image range are removed.
[0128] Triangulation is performed on each projection point on the image plane to obtain multiple triangular patches on the image plane.
[0129] For any triangular patch in the depth map, assuming the triangular patch is △ABC, the depth values dA, dB, and dC of 3D points A, B, and C can be obtained by projecting them onto the image plane.
[0130] For any pixel P within triangle ABC, the depth value of pixel P can be calculated as follows: .
[0131] Where S△PBC is the area of △PBC, S△PAC is the area of △PAC, S△PAB is the area of △PAB, and S△ABC is the area of △ABC.
[0132] The same applies to every other pixel within triangle ABC.
[0133] The same operation is performed for each other triangular facet in the depth map.
[0134] In this process, the rendered depth information minimizes the impact of noise, thereby improving the accuracy of the depth information in each of the resulting two-dimensional images.
[0135] In addition, the internal parameters of the handheld terminal include the focal length of the handheld terminal, which can remain unchanged during the process of capturing a two-dimensional image at this shooting position.
[0136] In one example, the homography matrix between two adjacent two-dimensional images captured at the shooting location can be obtained, and then the focal length of the handheld terminal can be decomposed from the homography matrix. For specific decomposition methods, please refer to existing methods, which will not be detailed here.
[0137] Secondly, the pose information of the handheld terminal when capturing each two-dimensional image at the shooting position can be determined based on the rotation angle of the stabilizer when capturing each two-dimensional image. For example, assuming that three two-dimensional images are captured sequentially at the shooting position, namely two-dimensional image A, two-dimensional image B, and two-dimensional image C, the rotation angle of the stabilizer is 0° when capturing two-dimensional image A, 45° when capturing two-dimensional image B, and 90° when capturing two-dimensional image C.
[0138] The difference between a rotation angle of 45° and a rotation angle of 0° is 45°, and the difference between a rotation angle of 90° and a rotation angle of 45° is also 45°. Therefore, 45° can be used as the difference information of the pose of the handheld terminal when capturing various two-dimensional images at this shooting position.
[0139] After capturing a two-dimensional image at the shooting location, the photographer can hold a handheld terminal with a distance sensor and stand at other shooting locations, and perform the above process at other shooting locations, so that at least two two-dimensional images and depth information of each two-dimensional image can be obtained at other shooting locations, until two-dimensional images are captured and depth information of the two-dimensional images is obtained at each shooting location.
[0140] Using the above method, a 360° x 180° image can be captured by rotating the handheld terminal (all-around image capture, including the entire spherical surface), which can then generate a 360° x 180° panoramic image.
[0141] In addition, the above method only requires a handheld terminal with a distance sensor, so the hardware cost of the above method is low.
[0142] Furthermore, in another embodiment of this application, for at least two two-dimensional images captured at any shooting location, the at least two two-dimensional images can be stitched together using a back-projection method (overlapping areas can be merged during the stitching process) to obtain a panoramic image of the three-dimensional space based on the shooting location. For example, if the at least two two-dimensional images are captured during a pure rotational movement of a handheld terminal at the shooting location, each pixel in the at least two two-dimensional images can be projected onto the corresponding latitude and longitude map according to the viewing direction to obtain a panoramic image of the three-dimensional space based on the shooting location.
[0143] Alternatively, at least two venue images can be stitched together using optical flow (overlapping areas can be merged during the stitching process) to obtain a panoramic view of the exhibition venue based on the acquisition location (this panoramic view obtained using optical flow has a low degree of imperfection).
[0144] The same applies to at least two two-dimensional images taken at each of the other shooting locations, thus obtaining a panoramic view of the three-dimensional space based on each shooting location.
[0145] Then, panoramic views of the three-dimensional space from each shooting location can be displayed for exhibitors to preview.
[0146] The back projection method yields panoramic images at a high rate, allowing the photographer to preview the panoramic image in a timely manner and determine whether the content in the panoramic image meets the requirements. If not, the image can be reshot promptly.
[0147] In one embodiment of this application, for the shooting mode of a handheld terminal without a distance sensor mounted on a stabilizer, the shooting strategy adapted to the shooting mode and the shooting device corresponding to the shooting mode capture a two-dimensional image of three-dimensional space and the depth information of the two-dimensional image, as detailed below: The handheld terminal can have an application installed, which may include an application corresponding to the stabilizer. The application can act as a communication interface between the handheld terminal and the stabilizer. That is, by operating the application on the handheld terminal, data interaction can be achieved between the handheld terminal and the stabilizer, such as controlling the operation of the stabilizer. Controlling the operation of the stabilizer includes controlling the rotation of the stabilizer (including rotation along the horizontal direction and rotation along the vertical direction). The handheld terminal and the stabilizer can communicate and connect based on Bluetooth, WiFi or NFC.
[0148] The handheld terminal has a shooting function, for example, the handheld terminal has a camera, which enables the handheld terminal to take pictures.
[0149] Stabilizers include two-axis stabilizers, three-axis stabilizers, four-axis stabilizers, or five-axis stabilizers, etc.
[0150] Furthermore, the filming crew can handhold the stabilizer or place it on a tripod, etc.
[0151] Thus, in this application, the photographer can manually select multiple shooting positions in three-dimensional space. The distance between each shooting position can be 5 meters, 6 meters, or 7 meters, or other values. This application does not limit the distance between each shooting position.
[0152] The photographer can place a handheld terminal without a distance sensor mounted on a stabilizer at one of the shooting positions, and then control the stabilizer to rotate (including horizontal and vertical rotation) through an application in the handheld terminal. During the rotation of the stabilizer, the photographer controls the handheld terminal camera to continuously capture images. The captured images include at least two frames, each of which includes a three-dimensional image. Thus, the at least two frames can be regarded as two-dimensional images in three-dimensional space.
[0153] Secondly, multiple 2D images can be stitched together to form a panoramic image. For example, multiple 2D images can be stitched together using back projection (overlapping areas can be merged during the stitching process) to obtain a panoramic image. Depth estimation is then performed on the panoramic image to obtain its depth information. This depth information is then projected onto the corresponding viewpoints of each 2D image to obtain the depth information of each individual 2D image.
[0154] In addition, the internal parameters of the handheld terminal include the focal length of the handheld terminal, which can remain unchanged during the process of capturing a two-dimensional image at this shooting position.
[0155] In one example, the homography matrix between two adjacent two-dimensional images captured at the shooting location can be obtained, and then the focal length of the handheld terminal can be decomposed from the homography matrix. For specific decomposition methods, please refer to existing methods, which will not be detailed here.
[0156] Secondly, the pose information of the handheld terminal when capturing each two-dimensional image at the shooting position can be determined based on the rotation angle of the stabilizer when capturing each two-dimensional image. For example, assuming that three two-dimensional images are captured sequentially at the shooting position, namely two-dimensional image A, two-dimensional image B, and two-dimensional image C, the rotation angle of the stabilizer is 0° when capturing two-dimensional image A, 45° when capturing two-dimensional image B, and 90° when capturing two-dimensional image C.
[0157] The difference between a rotation angle of 45° and a rotation angle of 0° is 45°, and the difference between a rotation angle of 90° and a rotation angle of 45° is also 45°. Therefore, 45° can be used as the difference information of the pose of the handheld terminal when capturing various two-dimensional images at this shooting position.
[0158] After capturing a 2D image at the shooting location, a handheld terminal without a distance sensor mounted on the stabilizer can be placed at other shooting locations, and the above process can be performed at other shooting locations to obtain at least two 2D images and depth information of each 2D image at other shooting locations, until 2D images are captured and depth information of the 2D images is obtained at each shooting location.
[0159] By combining the handheld terminal and the stabilizer in the above manner for joint control, the stabilizer's rotational movement is used to capture 360° x 180° images (all-around image capture, including the entire spherical surface). Due to the participation of the stabilizer, the influence of human shaking on the captured image can be avoided during the image capture process, and the center of the image capturing terminal can be located on the rotation axis of the stabilizer, so that the 360° x 180° panoramic image with the lowest possible level of defects or even no defects can be generated.
[0160] In addition, the above method only requires a stabilizer and a handheld terminal with a distance sensor, so the hardware cost of the above method is low.
[0161] Furthermore, in another embodiment of this application, for at least two two-dimensional images captured at any shooting location, the at least two two-dimensional images can be stitched together using a back-projection method (overlapping areas can be merged during the stitching process) to obtain a panoramic image of the three-dimensional space based on the shooting location. For example, if the at least two two-dimensional images are captured during a pure rotational movement of a handheld terminal at the shooting location, each pixel in the at least two two-dimensional images can be projected onto the corresponding latitude and longitude map according to the viewing direction to obtain a panoramic image of the three-dimensional space based on the shooting location.
[0162] Alternatively, at least two venue images can be stitched together using optical flow (overlapping areas can be merged during the stitching process) to obtain a panoramic view of the exhibition venue based on the acquisition location (this panoramic view obtained using optical flow has a low degree of imperfection).
[0163] The same applies to at least two two-dimensional images taken at each of the other shooting locations, thus obtaining a panoramic view of the three-dimensional space based on each shooting location.
[0164] Then, panoramic views of the three-dimensional space from each shooting location can be displayed for exhibitors to preview.
[0165] The back projection method yields panoramic images at a high rate, allowing the photographer to preview the panoramic image in a timely manner and determine whether the content in the panoramic image meets the requirements. If not, the image can be reshot promptly.
[0166] In one embodiment of this application, for the shooting mode of a handheld terminal without a distance sensor, a two-dimensional image of three-dimensional space and depth information of the two-dimensional image are captured according to the shooting strategy adapted to the shooting mode and the shooting device of the physical form corresponding to the shooting mode, as follows: The handheld terminal has a shooting function, for example, the handheld terminal has a camera, which enables the handheld terminal to take pictures.
[0167] Thus, in this application, the photographer can manually select multiple shooting positions in three-dimensional space. The distance between each shooting position can be 5 meters, 6 meters, or 7 meters, or other values. This application does not limit the distance between each shooting position.
[0168] The handheld terminal can have an application installed for taking pictures. The person taking the picture can hold the handheld terminal with a distance sensor and stand at one of the shooting positions to take pictures.
[0169] During the shooting process, the handheld terminal can generate shooting guidance information (such as a circle) and a shooting crosshair (such as a ring) on the screen based on the data from the gyroscope in the handheld terminal. The user can move the position of the handheld terminal in space and change the orientation of the handheld terminal's camera to adjust the shooting crosshair so that the shooting crosshair coincides with the shooting guidance information. Then, the user can control the camera to capture a frame of image and then rotate the handheld terminal horizontally.
[0170] During the rotation of the handheld terminal, the handheld terminal can generate another shooting guidance information and shooting crosshair on the screen based on the data from the gyroscope in the handheld terminal. The user can move the position of the handheld terminal in space and change the orientation of the handheld terminal's camera to adjust the shooting crosshair so that the shooting crosshair coincides with the shooting guidance information. Then, the user can control the camera to capture one frame of image. In this way, the captured image includes at least two frames of image, and each of the at least two frames of image includes a three-dimensional space scene. Thus, each of the at least two frames of image can be regarded as a two-dimensional image in three-dimensional space.
[0171] Secondly, multiple 2D images can be stitched together to form a panoramic image. For example, multiple 2D images can be stitched together using back projection (overlapping areas can be merged during the stitching process) to obtain a panoramic image. Depth estimation is then performed on the panoramic image to obtain its depth information. This depth information is then projected onto the corresponding viewpoints of each 2D image to obtain the depth information of each individual 2D image.
[0172] In addition, the internal parameters of the handheld terminal include its focal length, which can remain constant during the capture of a two-dimensional image at this shooting position. In one example, the homography matrix between two adjacent two-dimensional images captured at this shooting position can be obtained, and then the focal length of the handheld terminal can be decomposed from the homography matrix. For specific decomposition methods, please refer to existing methods, which will not be detailed here.
[0173] Secondly, the pose information of the handheld terminal when capturing each 2D image at this shooting position can be determined based on the rotation angle of the stabilizer when capturing each 2D image. For example, assuming three 2D images, A, B, and C, are captured sequentially at this shooting position, the stabilizer's rotation angle is 0° when capturing 2D image A, 45° when capturing 2D image B, and 90° when capturing 2D image C. The difference between a rotation angle of 45° and a rotation angle of 0° is 45°, and the difference between a rotation angle of 90° and a rotation angle of 45° is also 45°. Therefore, 45° can be used as the difference information of the handheld terminal's pose information when capturing each 2D image at this shooting position.
[0174] After capturing a two-dimensional image at the shooting location, the photographer can hold a handheld terminal without a distance sensor and stand at other shooting locations, and perform the above process at other shooting locations, so that at least two two-dimensional images and depth information of each two-dimensional image can be obtained at other shooting locations, until two-dimensional images are captured and depth information of the two-dimensional images are obtained at each shooting location.
[0175] Using the above method, a 360° x 180° image can be captured by rotating the handheld terminal (all-around image capture, including the entire spherical surface), which can then generate a 360° x 180° panoramic image.
[0176] In addition, the above method only requires a handheld terminal without a distance sensor, so the hardware cost of the above method is low.
[0177] Furthermore, in another embodiment of this application, for at least two two-dimensional images captured at any shooting location, the at least two two-dimensional images can be stitched together using a back-projection method (overlapping areas can be merged during the stitching process) to obtain a panoramic view of the three-dimensional space based on that shooting location. For example, if at least two two-dimensional images are captured during a pure rotational movement of a handheld terminal at that shooting location, each pixel in the at least two two-dimensional images can be projected onto the corresponding latitude and longitude map according to the viewing direction to obtain a panoramic view of the three-dimensional space based on that shooting location. Alternatively, at least two venue images can be stitched together using an optical flow method (overlapping areas can be merged during the stitching process) to obtain a panoramic view of the exhibition venue based on that acquisition location (this panoramic view obtained using the optical flow method has a low degree of imperfection). The same applies to at least two two-dimensional images captured at each other shooting location, thereby obtaining panoramic views of the three-dimensional space based on each shooting location. The panoramic views of the three-dimensional space at each shooting location can then be displayed for exhibitors to preview. The back projection method yields panoramic images at a high rate, allowing the photographer to preview the panoramic image in a timely manner and determine whether the content in the panoramic image meets the requirements. If not, the image can be reshot promptly.
[0178] In one embodiment of this application, for a shooting mode combining a handheld terminal without a distance sensor and a panoramic camera, the shooting strategy adapted to the shooting mode and the shooting device corresponding to the shooting mode capture a two-dimensional image of the three-dimensional space and the depth information of the two-dimensional image, as detailed below: The handheld terminal can have an application installed, which may include an application corresponding to the panoramic camera. This application acts as a communication interface between the handheld terminal and the panoramic camera; that is, by operating the application on the handheld terminal, data interaction between the two cameras can be achieved, such as controlling the panoramic camera to take pictures. The handheld terminal and the panoramic camera can communicate and connect via Bluetooth, WiFi, or NFC. Furthermore, the photographer can place a stabilizer on a tripod, etc. Thus, in this application, the photographer can manually select multiple shooting positions in three-dimensional space. The distance between these shooting positions can include 5 meters, 6 meters, or 7 meters, or other values; this application does not limit the distance between the shooting positions.
[0179] The photographer can place the panoramic camera at one of the shooting positions and then control the panoramic camera to capture images through an application in a handheld terminal. The captured images include a two-dimensional panoramic image, and each two-dimensional panoramic image includes a scene in three-dimensional space. Thus, a two-dimensional panoramic image can be regarded as a two-dimensional image in three-dimensional space.
[0180] Depth estimation can be performed on the panoramic image to obtain its depth information. Then, the depth information of the panoramic image can be projected onto the corresponding viewpoints of each two-dimensional image to obtain the depth information of each two-dimensional image.
[0181] In addition, the internal parameters of the handheld terminal include its focal length, which can remain constant during the capture of a two-dimensional image at this shooting position. In one example, the homography matrix between two adjacent two-dimensional images captured at this shooting position can be obtained, and then the focal length of the handheld terminal can be decomposed from the homography matrix. For specific decomposition methods, please refer to existing methods, which will not be detailed here.
[0182] Secondly, the pose information of the handheld terminal when capturing each two-dimensional image at the shooting position can be determined based on the rotation angle of the stabilizer when capturing each two-dimensional image. For example, assuming that three two-dimensional images are captured sequentially at the shooting position, namely two-dimensional image A, two-dimensional image B, and two-dimensional image C, the rotation angle of the stabilizer is 0° when capturing two-dimensional image A, 45° when capturing two-dimensional image B, and 90° when capturing two-dimensional image C.
[0183] The difference between a rotation angle of 45° and a rotation angle of 0° is 45°, and the difference between a rotation angle of 90° and a rotation angle of 45° is also 45°. Therefore, 45° can be used as the difference in pose information of the handheld terminal when capturing various two-dimensional images at this shooting position. By combining the handheld terminal and the panoramic camera together for joint control in the above manner, and using the panoramic camera to capture 360° x 180° images (omnidirectional image capture, including the entire spherical surface), a 360° x 180° panoramic image can be directly obtained.
[0184] In addition, the above method only requires a handheld terminal without a distance sensor, a panoramic camera, and a handheld terminal with a distance sensor. Therefore, the hardware cost of the above method is low.
[0185] Furthermore, in another embodiment of this application, for at least two two-dimensional images captured at any shooting location, the at least two two-dimensional images can be stitched together using a back-projection method (overlapping areas can be merged during the stitching process) to obtain a panoramic image of the three-dimensional space based on the shooting location. For example, if the at least two two-dimensional images are captured during a pure rotational movement of a handheld terminal at the shooting location, each pixel in the at least two two-dimensional images can be projected onto the corresponding latitude and longitude map according to the viewing direction to obtain a panoramic image of the three-dimensional space based on the shooting location.
[0186] Alternatively, at least two venue images can be stitched together using optical flow (overlapping areas can be merged during the stitching process) to obtain a panoramic view of the exhibition venue based on the acquisition location (this panoramic view obtained using optical flow has a low degree of imperfection).
[0187] The same applies to at least two two-dimensional images taken at each of the other shooting locations, thus obtaining a panoramic view of the three-dimensional space based on each shooting location.
[0188] Then, panoramic views of the three-dimensional space from each shooting location can be displayed for exhibitors to preview.
[0189] The back projection method yields panoramic images at a high rate, allowing the photographer to preview the panoramic image in a timely manner and determine whether the content in the panoramic image meets the requirements. If not, the image can be reshot promptly.
[0190] In this application, multiple two-dimensional images and their depth information are obtained from multiple shooting positions in three-dimensional space. Then, the content in the two-dimensional images taken at each shooting position can be identified, and the similarity between the content in the two-dimensional images taken at each shooting position can be analyzed. For example, for any shooting position, multiple two-dimensional images taken at the shooting position are stitched together to obtain a panoramic image, and the panoramic image corresponding to that shooting position is obtained. The same process is repeated for each other shooting position to obtain a panoramic image corresponding to each shooting position. Then, the similarity between the panoramic images corresponding to each shooting position is analyzed.
[0191] Based on the similarity between panoramic images corresponding to each shooting location and the relative positional relationships between these locations, topological relationships are generated. These topological relationships include multiple shooting locations and their interdependencies; they are then used to generate a 3D virtual space model. For example, a 3D virtual space model can be generated based on 2D images of 3D space taken from multiple shooting locations and the topological relationships.
[0192] The interdependence between multiple shooting locations can be understood as follows: any one shooting location is interdependent with at least one other shooting location. For example, if one shooting location depends on another, it means that in the process of generating a 3D virtual space model, a 2D image taken at one shooting location can be stitched together with a 2D image taken at the other shooting location.
[0193] Topological relationships can be represented in the form of a directed graph. For example, if a certain shooting position depends on another shooting position, then in the topological relationship, an arrow is used to point from the other shooting position to the first shooting position.
[0194] Secondly, it is possible to obtain a top-down map (directed graph) of three-dimensional space, mark each shooting position on the top-down map, and indicate the interdependence between each shooting position with arrows, so that the shooting party can check whether there are redundant shooting positions or whether any shooting positions have been missed by using the top-down map.
[0195] If there are redundant shooting locations, the 2D images taken at those locations, as well as the depth information of those 2D images, can be deleted.
[0196] If a shooting location is missed, a 2D image can be captured at the missed shooting location, and the depth information of the 2D image can be obtained. The missed shooting location and the interdependence between the missed shooting location and other shooting locations can be added to the topological relationship.
[0197] In this application, the internal parameters of the handheld terminal, the difference information of pose information corresponding to each acquisition position, and the panoramic images corresponding to each shooting position (obtained respectively based on the two-dimensional images (and / or the depth information of the two-dimensional images) captured at each shooting position) can be obtained. The internal parameters of the handheld terminal, the difference information of pose information corresponding to each acquisition position, the two-dimensional images captured at each acquisition position, and the depth information of the two-dimensional images can be processed according to the SFM algorithm (Structure From Motion, which is a three-dimensional reconstruction method used to achieve three-dimensional reconstruction from camera motion) to complete the three-dimensional reconstruction of the three-dimensional space and obtain the initial three-dimensional virtual space model. Then, the panoramic images corresponding to each acquisition position and the initial three-dimensional virtual space model are fused with TM (Texture Mapping) technology to obtain the final three-dimensional virtual space model.
[0198] When at least one shooting mode is displayed on the screen of the handheld terminal, it can be determined whether the handheld terminal has a distance sensor and whether the handheld terminal is mounted on a stabilizer; if the handheld terminal has a distance sensor and is mounted on a stabilizer, at least the shooting mode of the handheld terminal with a distance sensor mounted on the stabilizer is displayed on the screen of the handheld terminal; or, if the handheld terminal has a distance sensor but is not mounted on a stabilizer, at least the shooting mode of the handheld terminal with a distance sensor is displayed on the screen of the handheld terminal; or, if the handheld terminal does not have a distance sensor but is mounted on a stabilizer, at least the shooting mode of the handheld terminal without a distance sensor mounted on the stabilizer is displayed on the screen of the handheld terminal; or, if the handheld terminal does not have a distance sensor and is not mounted on a stabilizer, at least the shooting mode of the handheld terminal without a distance sensor and / or the shooting mode of the handheld terminal without a distance sensor combined with a panoramic camera is displayed on the screen of the handheld terminal.
[0199] Specifically, when at least the handheld terminal's screen displays a shooting mode for a handheld terminal without a distance sensor and / or a shooting mode combining a handheld terminal without a distance sensor and a panoramic camera, it can be determined whether the handheld terminal has a communication connection with the panoramic camera; if there is no communication connection between the handheld terminal and the panoramic camera, at least the handheld terminal's screen displays a shooting mode for a handheld terminal without a distance sensor; or, if there is a communication connection between the handheld terminal and the panoramic camera, at least the handheld terminal's screen displays a shooting mode for a handheld terminal without a distance sensor and / or a shooting mode combining a handheld terminal without a distance sensor and a panoramic camera.
[0200] For example, a shooting mode of a handheld terminal without a distance sensor is displayed on the screen of the handheld terminal in a first display style, and a shooting mode of a handheld terminal without a distance sensor combined with a panoramic camera is displayed on the screen of the handheld terminal in a second display style; wherein the salience of the first display style is less than the salience of the second display style.
[0201] Regarding salience, the easier it is for users to see the shooting mode on the screen, the higher the salience of the shooting mode's display method; conversely, the harder it is for users to see the shooting mode on the screen, the lower its salience. In one example, the display method includes its display position on the screen and its display style. The display size of the display style included in the second display method is larger than the display size of the display style included in the first display method. The display position of the display style included in the second display method is closer to the center of the screen, while the display position of the display style included in the first display method is closer to the edge of the screen, etc.
[0202] Specifically, when determining whether a handheld terminal has a distance sensor and whether it is mounted on a stabilizer, configuration options for the distance sensor and the stabilizer can be displayed on the screen. If both the distance sensor configuration option and the stabilizer configuration option are selected, it is determined that the handheld terminal has a distance sensor and is mounted on a stabilizer. Alternatively, if both the distance sensor configuration option and the stabilizer configuration option are not selected, it is determined that the handheld terminal has a distance sensor but is not mounted on a stabilizer. Alternatively, if both the distance sensor configuration option and the stabilizer configuration option are selected, it is determined that the handheld terminal does not have a distance sensor but is mounted on a stabilizer. Alternatively, if both the distance sensor configuration option and the stabilizer configuration option are not selected, it is determined that the handheld terminal does not have a distance sensor and is not mounted on a stabilizer.
[0203] When determining whether a handheld terminal is mounted on a stabilizer, it can be determined whether the handheld terminal is in communication with the stabilizer; if the handheld terminal is in communication with the stabilizer, it is determined that the handheld terminal is mounted on the stabilizer; or, if the handheld terminal is not in communication with the stabilizer, it is determined that the handheld terminal is not mounted on the stabilizer.
[0204] Reference Figure 4 The flowchart illustrates the steps of a data processing method according to this application, which is applied to... Figure 1 In one of the shooting terminals 02 or cloud 01 shown, the method may specifically include the following steps: In step S301, a spherical projection panoramic image is generated based on multiple two-dimensional images; and a hexahedral projection panoramic image is generated based on multiple two-dimensional images.
[0205] Multiple two-dimensional images can be captured at a single shooting position, such as by using a shooting device to rotate horizontally at that shooting position.
[0206] In step S302, the depth information of the spherical projection panoramic image is obtained based on the spherical projection panoramic image, the hexahedral projection panoramic image, and the depth information acquisition model.
[0207] In this application, a depth information acquisition model can be trained in advance. The training process includes: acquiring multiple sample datasets, each of which includes a sample spherical projection panoramic image generated from multiple two-dimensional images, a sample hexahedral projection panoramic image generated from multiple two-dimensional images, and standard depth information of the sample spherical projection panoramic image.
[0208] Then, the model is trained on multiple sample datasets until the parameters in the model converge, thus obtaining a deep information acquisition model. This model includes models such as the U-net model.
[0209] Thus, in this step, the spherical projection panoramic image and the hexahedral projection panoramic image can be input into the depth information acquisition model to obtain the depth information of the spherical projection panoramic image output by the depth information recognition model.
[0210] Among them, the spherical projection panoramic image generated from multiple two-dimensional images often contains distortions, and the degree of distortion is usually greater closer to the top and bottom of the sphere. If the depth information of the spherical projection panoramic image is obtained solely from the spherical projection panoramic image, the depth information of the distorted parts of the spherical projection panoramic image is often inaccurate.
[0211] In this application, depth information of the spherical projection panoramic image is obtained by combining a spherical projection panoramic image and a hexahedral projection panoramic image. The hexahedral projection panoramic image is also generated from multiple two-dimensional images. The hexahedral projection panoramic image has lower distortion; for example, the distortion at the top and bottom of the hexahedron is typically low. Therefore, the hexahedral projection panoramic image is used to compensate for the distortion problem of the spherical projection panoramic image, thereby improving the accuracy of the obtained depth information. For example, it can at least improve the accuracy of the depth information of the distorted parts of the spherical projection panoramic image.
[0212] In one embodiment, the depth information acquisition model is an improvement on the Unet model, which can employ a fully convolutional neural network. The depth information acquisition model includes a spherical encoding network, a hexahedral encoding network, and a decoding network.
[0213] A spherical encoding network is used to upsample the spherical projection panorama, obtaining spherical image features at different scales. A hexahedral encoding network is used to upsample the hexahedral projection panorama, obtaining hexahedral image features at different scales. A decoding network is used to downsample the spherical image features and hexahedral image features at different scales, obtaining the depth information of the spherical projection panorama.
[0214] See Figure 5 A spherical coding network can include multiple cascaded spherical coding layers. A hexahedral coding network can include multiple cascaded hexahedral coding layers. A decoding network can include multiple cascaded decoding layers. The number of spherical coding layers, the number of hexahedral coding layers, and the number of decoding layers can be the same.
[0215] The spherical projection panoramic image can be input into the first cascaded spherical coding layer, which encodes the spherical projection panoramic image to obtain the first spherical image feature. Then, the first spherical image feature can be input into the second cascaded spherical coding layer, which encodes the first spherical image feature to obtain the second spherical image feature, and so on. The (N-1)th spherical image feature can be input into the Nth cascaded spherical coding layer, which processes the (N-1)th spherical image feature to obtain the Nth spherical image feature.
[0216] The hexahedral projection panoramic image can be input into the first concatenated hexahedral encoding layer. The first hexahedral encoding layer encodes the hexahedral projection panoramic image to obtain the first hexahedral image feature. Similarly, the first hexahedral image feature can be input into the second concatenated hexahedral encoding layer, which encodes the first hexahedral image feature to obtain the second hexahedral image feature, and so on. The (N-1)th hexahedral image feature can be input into the Nth concatenated hexahedral encoding layer, which processes the (N-1)th hexahedral image feature to obtain the Nth hexahedral image feature.
[0217] N is greater than or equal to 2. N represents the number of spherical coding layers, the number of hexahedral coding layers, and the number of decoding layers.
[0218] In addition to multiple cascaded decoding layers, the decoding network may also include multiple fusion layers. The number of fusion layers, the number of spherical coding layers, the number of hexahedral coding layers, and the number of decoding layers can be the same.
[0219] The first spherical image feature and the first hexahedral image feature can be input into the first fusion layer. The first fusion layer can fuse the first spherical image feature and the first hexahedral image feature to obtain the first fused image feature. The second spherical image feature and the second hexahedral image feature can be input into the second fusion layer. The second fusion layer can fuse the second spherical image feature and the second hexahedral image feature to obtain the second fused image feature, and so on. The Nth spherical image feature and the Nth hexahedral image feature can be input into the Nth fusion layer. The Nth fusion layer can fuse the Nth spherical image feature and the Nth hexahedral image feature to obtain the Nth fused image feature.
[0220] The Nth fused image feature can be input into the Nth decoder layer in a series. The Nth decoder can process the Nth fused image feature to obtain the Nth depth feature. The (N-1)th fused image feature and the Nth depth feature can be input into the (N-1)th decoder layer in a series. The (N-1)th decoder layer can process the (N-1)th fused image feature and the Nth depth feature to obtain the (N-1)th depth feature, and so on, until the 1st fused image feature and the 2nd depth feature are input into the 1st decoder layer in a series. The 1st decoder layer can process the 1st fused image feature and the 2nd depth feature to obtain the 1st depth feature, which is used as the depth information of the spherical projection panoramic image.
[0221] In one embodiment, see Figure 6 For the fusion layer, it can include: first Cat (Concatenate) layer, second Cat layer, C2E layer, hybrid layer, first 1*1 convolutional kernel, first BN (Batch-Normalization) layer, first ReLU (Rectified Linear Unit) function, 3*3 convolutional kernel, second BN layer, SE (Squeeze-and-Excitation) layer, second 1*1 convolutional kernel, and second ReLU, etc.
[0222] The input terminals of the first Cat layer and the second Cat layer are used to input spherical image features. The input terminal of the C2E layer is used to input hexahedral image features. The output terminal of the C2E layer is connected to the input terminal of the first Cat layer and the input terminal of the mixing layer. The output terminal of the first Cat layer is connected to the input terminal of the first 1*1 convolutional kernel. The output terminal of the first 1*1 convolutional kernel is connected to the input terminal of the first BN layer. The output terminal of the first BN layer is connected to the input terminal of the first ReLU function. The output terminal of the first ReLU function is connected to the input terminal of the 3*3 convolutional kernel. The output terminal of the 3*3 convolutional kernel is connected to the input terminal of the second BN layer. The output terminal of the second BN layer is connected to the input terminal of the mixing layer. The output terminal of the mixing layer is connected to the input terminal of the second Cat layer. The output terminal of the second Cat layer is connected to the input terminal of the SE layer. The output terminal of the SE layer is connected to the input terminal of the second 1*1 convolutional kernel. The output terminal of the second 1*1 convolutional kernel is connected to the input terminal of the second ReLU function. The output terminal of the second ReLU function is the output terminal of the fusion layer, used to output fused image features.
[0223] Among them, a spherical projection panorama is obtained by projecting multiple two-dimensional images onto a spherical coordinate system, while a hexahedral projection panorama is obtained by projecting multiple two-dimensional images onto a hexahedral coordinate system. Therefore, spherical image features are image features corresponding to the spherical coordinate system, and hexahedral image features are image features corresponding to the hexahedral coordinate system. The C2E layer is used to convert hexahedral image features into image features corresponding to the spherical coordinate system.
[0224] In particular, hexahedral projection panoramic images often contain image discontinuities, which can affect the accuracy of depth information in the obtained spherical projection panoramic images. Therefore, it is necessary to minimize the image discontinuities that often exist in hexahedral projection panoramic images.
[0225] To mitigate the discontinuities often present in hexahedral projection panoramic images, this application employs a first 1x1 convolution kernel, a first batch normalization (BN) layer, a first ReLU function, a 3x3 convolution kernel, a second BN layer, and a hybrid layer to perform residual modulation on the process of converting hexahedral image features into corresponding spherical coordinate system image features. This approach can alleviate the discontinuities often present in hexahedral projection panoramic images, thereby improving the accuracy of the depth information obtained from the spherical projection panoramic image.
[0226] Furthermore, in this application, the encoding network includes two encoding networks: a spherical encoding network and a hexahedral encoding network. The image features output by the spherical encoding network and the hexahedral encoding network are fused together and input into the decoding network. The decoding network consists of only one decoding network, which does not require feature fusion. Therefore, the depth information acquisition model in this application only involves the fusion of the encoding networks, not the fusion of the decoding networks. Through multiple comparative experiments, it has been found that compared to a multi-path decoding network that involves feature fusion, the single decoding network in this application, which does not involve feature fusion, improves the accuracy of the depth information of the final acquired spherical projection panoramic image, reduces the number of network parameters required, lowers the computational resources required, and reduces computation time.
[0227] The SE layer includes a global pooling layer, a first fully connected layer, a third ReLU function, a second fully connected layer, and a sigmoid function. The input of the global pooling layer is connected to the output of the second CAT layer. The output of the first fully connected layer is connected to the input of the third ReLU function. The output of the third ReLU function is connected to the input of the second fully connected layer. The output of the second fully connected layer is connected to the input of the sigmoid function. The output of the sigmoid function is the output of the SE layer, which is connected to the input of the second 1x1 convolutional kernel. This fusion layer, through which spherical and hexahedral image features are fused more accurately and completely, allows for a more comprehensive and accurate fusion of these features.
[0228] In one embodiment, for the fusion layer, see [link to fusion layer]. Figure 7This can include: a first 3x3 convolutional kernel, a first ReLU function, a Cat layer, a C2E layer, a second 3x3 convolutional kernel, a second ReLU function, a 1x1 convolutional kernel, a Sigmoid function, a Mask layer, a multiplication layer, and a hybrid layer. The input of the first 3x3 convolutional kernel is used to input spherical image features. The input of the C2E layer is used to input hexahedral image features. The output of the first 3x3 convolutional kernel is connected to the input of the first ReLU function. The output of the first ReLU function is connected to the input of the Cat layer. The output of the C2E layer is connected to the input of the second 3x3 convolutional kernel. The output of the second 3x3 convolutional kernel is connected to the input of the second ReLU function. The output of the second ReLU function is connected to the input of the Cat layer. The output of the Cat layer is connected to the input of the 1x1 convolutional kernel. The input of the 1x1 convolutional kernel is connected to the input of the Sigmoid function. The output of the Sigmoid function is connected to the input of the multiplication layer. The output of the C2E layer is also connected to the input of the multiplication layer. The output of the multiplication layer (including the Wise Prod layer, etc.) connects to the input of the mixing layer. The input of the mixing layer is also used to input spherical image features. The output of the mixing layer is the output of the fusion layer, used to output fused image features.
[0229] In one embodiment, for the fusion layer, see [link to fusion layer]. Figure 8 It can include: Cat layer, C2E layer, 1*1 convolution kernel and ReLU function, etc.
[0230] The input to the Cat layer is used to input spherical image features. The input to the C2E layer is used to input hexahedral image features. The output of the C2E layer is connected to the input of the Cat layer. The output of the Cat layer is connected to the input of a 1x1 convolutional kernel. The output of the 1x1 convolutional kernel is connected to the input of the ReLU function. The output of the ReLU function is the output of the fusion layer, used to output fused image features.
[0231] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions involved are not necessarily required by this application.
[0232] Reference Figure 9The diagram illustrates a structural block diagram of a data processing apparatus according to this application, applied to a shooting terminal. The apparatus may include the following modules: a first acquisition module 11, used to acquire the original topological relationship, which includes: multiple shooting positions in the three-dimensional space involved in the previous establishment of the three-dimensional virtual space model, and the interdependencies between the multiple shooting positions; a first determination module 12, used to determine the target shooting position among the multiple shooting positions in the original topological relationship, which is the target shooting position for which the shooting content needs to be changed; a second acquisition module 13, used to acquire the target interdependency between the target shooting position and the shooting positions in the original topological relationship based on the original topological relationship and the target shooting position; and to acquire a new two-dimensional image of the three-dimensional space captured at the target shooting position; and a sending module 14, used to send the target shooting position, the target interdependency, and the new two-dimensional image to the cloud, so that the cloud generates a new three-dimensional virtual space model of the three-dimensional space based at least on the target shooting position, the target interdependency, the new two-dimensional image, and the old two-dimensional image; wherein the old two-dimensional image includes: two-dimensional images captured at shooting positions other than the target shooting position among the multiple shooting positions involved in the previous establishment of the three-dimensional virtual space model.
[0233] Through this application, if the 3D virtual space model needs to adapt to the changed parts of the 3D space due to changes in some scenes arranged in the 3D space, the shooting party does not need to use the 2D images of the 3D space taken by the shooting party terminal at each required shooting position in the 3D space again, the shooting party terminal does not need to acquire the 2D images of the 3D space taken at each required shooting position in the 3D space again, and does not need to send the 2D images of the 3D space taken at each required shooting position in the 3D space to the cloud again.
[0234] The camera operator can use its terminal to capture new 2D images of the target location in 3D space, focusing solely on the content to be captured. The terminal can then send only these new 2D images to the cloud. The cloud can then use these new images, along with previously existing ones, to generate a new 3D virtual space model. This reduces the amount of data required for updates, thus lowering labor costs and conserving system and network resources for the camera operator. The more frequently the 3D space is updated, the more significant the savings in labor, system, and network resources become.
[0235] Reference Figure 10This diagram illustrates a structural block diagram of a data processing apparatus according to this application, applied in the cloud. The apparatus may include the following modules: a receiving module 21, used to receive a new two-dimensional image of a three-dimensional space captured at a target shooting location, sent by a shooting terminal; the target interdependency relationship between the target shooting location and shooting locations in the original topological relationship; the target interdependency relationship is obtained based on the original topological relationship and the target shooting location; the original topological relationship includes: multiple shooting locations in the three-dimensional space involved in the previous establishment of the three-dimensional virtual space model and the interdependencies between these multiple shooting locations; the target shooting location is the shooting location among the multiple shooting locations in the original topological relationship where the content to be shot needs to be changed; a third obtaining module 22, used to obtain cached old two-dimensional images, including: two-dimensional images captured at shooting locations other than the target shooting location among the multiple shooting locations involved in the previous establishment of the three-dimensional virtual space model; and a first generating module 23, used to generate a new three-dimensional virtual space model of the three-dimensional space based on the target shooting location, the target shooting dependency relationship, the new two-dimensional image, and the old two-dimensional image.
[0236] Through this application, if the 3D virtual space model needs to adapt to the changed parts of the 3D space due to changes in some scenes arranged in the 3D space, the shooting party does not need to use the 2D images of the 3D space taken by the shooting party terminal at each required shooting position in the 3D space again, the shooting party terminal does not need to acquire the 2D images of the 3D space taken at each required shooting position in the 3D space again, and does not need to send the 2D images of the 3D space taken at each required shooting position in the 3D space to the cloud again. The camera operator can use its terminal to capture new 2D images of the target location in 3D space, focusing solely on the content to be captured. The terminal can then send only these new 2D images to the cloud. The cloud can then use these new images, along with previously existing ones, to generate a new 3D virtual space model. This reduces the amount of data required for updates, thus lowering labor costs and conserving system and network resources for the camera operator. The more frequently the 3D space is updated, the more significant the savings in labor, system, and network resources become.
[0237] Reference Figure 11This diagram illustrates a structural block diagram of a data processing apparatus according to this application, applied to a handheld terminal. The apparatus may include the following modules: a display module 31, used to display at least one shooting mode on the screen of the handheld terminal; the shooting modes include at least: a shooting mode of a handheld terminal with a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal without a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal with a distance sensor, a shooting mode of a handheld terminal without a distance sensor, and a shooting mode of a handheld terminal without a distance sensor combined with a panoramic camera; a second determining module 32, used to determine the selected target shooting mode among the at least one displayed shooting modes; and a shooting module 33, used to capture a two-dimensional image of three-dimensional space and depth information of the two-dimensional image according to the target shooting strategy adapted to the target shooting mode and the shooting device corresponding to the physical form of the target shooting mode.
[0238] The shooting devices corresponding to the shooting modes include: handheld terminals, or a combination of handheld terminals and panoramic cameras, or a stabilizer. The handheld terminals support handheld terminals without distance sensors or handheld terminals with distance sensors.
[0239] This application enables the capture of 2D images in 3D space using only a handheld terminal, eliminating the need for professional shooting equipment. Due to the low cost and high ownership rate of handheld terminals, this application reduces shooting costs and increases convenience. When auxiliary accessories are needed to improve image quality, stabilizers and panoramic cameras are readily available and inexpensive. Compared to using expensive professional equipment, this application saves costs while maintaining the same or similar image quality.
[0240] Secondly, when the photographer has a handheld terminal, which may or may not have a distance sensor, and the photographer may or may not have a stabilizer, and may or may not have a panoramic camera, this application can support multiple shooting modes so that the photographer can choose the appropriate shooting mode when needed. For example, the photographer can choose the appropriate shooting mode based on the physical shooting equipment they currently have, thereby improving the flexibility of shooting.
[0241] Reference Figure 12The diagram shows a structural block diagram of a data processing apparatus according to this application. The apparatus may include the following modules: a second generation module 41, used to generate a spherical projection panoramic image based on multiple two-dimensional images; and to generate a hexahedral projection panoramic image based on multiple two-dimensional images; and a fourth acquisition module 42, used to acquire the depth information of the spherical projection panoramic image based on the spherical projection panoramic image, the hexahedral projection panoramic image, and a depth information acquisition model.
[0242] Among them, the spherical projection panoramic image generated from multiple two-dimensional images often contains distortions, and the degree of distortion is usually greater closer to the top and bottom of the sphere. If the depth information of the spherical projection panoramic image is obtained solely from the spherical projection panoramic image, the depth information of the distorted parts of the spherical projection panoramic image is often inaccurate.
[0243] In this application, depth information of the spherical projection panoramic image is obtained by combining a spherical projection panoramic image and a hexahedral projection panoramic image. The hexahedral projection panoramic image is also generated from multiple two-dimensional images. The hexahedral projection panoramic image has lower distortion; for example, the distortion at the top and bottom of the hexahedron is typically low. Therefore, the hexahedral projection panoramic image is used to compensate for the distortion problem of the spherical projection panoramic image, thereby improving the accuracy of the obtained depth information. For example, it can at least improve the accuracy of the depth information of the distorted parts of the spherical projection panoramic image.
[0244] This application also provides a non-volatile readable storage medium storing one or more modules (programs). When these modules are applied to a device, they enable the device to execute the instructions for the method steps in this application.
[0245] This application provides one or more machine-readable media storing instructions that, when executed by one or more processors, cause an electronic device to perform one or more methods as described in the above embodiments. In this application, the electronic device includes a server, a gateway, sub-devices, etc., and the sub-devices are devices such as Internet of Things (IoT) devices.
[0246] Embodiments of this disclosure can be implemented as an apparatus with any suitable hardware, firmware, software, or any combination thereof, configured as desired. This apparatus may include electronic devices such as servers (clusters) and terminal devices such as IoT devices.
[0247] Figure 13 An exemplary apparatus 1300 that can be used to implement the various embodiments of this application is schematically illustrated. For one embodiment, Figure 13An exemplary device 1300 is shown, which includes one or more processors 1302, a control module (chipset) 1304 coupled to at least one of the processors 1302, a memory 1306 coupled to the control module 1304, a non-volatile memory (NVM) / storage device 1308 coupled to the control module 1304, one or more input / output devices 1310 coupled to the control module 1304, and a network interface 1312 coupled to the control module 1304.
[0248] Processor 1302 may include one or more single-core or multi-core processors, and processor 1302 may include any combination of general-purpose processors or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, device 1300 can function as a server device such as a gateway in the embodiments of this application.
[0249] In some embodiments, apparatus 1300 may include one or more computer-readable media (e.g., memory 1306 or NVM / storage device 1308) having instructions 1314 and one or more processors 1302 that are combined with the one or more computer-readable media and configured to execute instructions 1314 to implement modules and thus perform the actions in this disclosure.
[0250] In one embodiment, the control module 1304 may include any suitable interface controller to provide any suitable interface to at least one of the processors 1302 and / or any suitable device or component communicating with the control module 1304.
[0251] The control module 1304 may include a memory controller module to provide an interface to the memory 1306. The memory controller module may be a hardware module, a software module, and / or a firmware module.
[0252] Memory 1306 may be used, for example, to load and store data and / or instructions 1314 for device 1300. In one embodiment, memory 1306 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, memory 1306 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).
[0253] In one embodiment, control module 1304 may include one or more input / output controllers to provide an interface to NVM / storage device 1308 and (one or more) input / output devices 1310. For example, NVM / storage device 1308 may be used to store data and / or instructions 1314. NVM / storage device 1308 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drive (HDD), one or more optical disc (CD) drives, and / or one or more digital universal optical disc (DVD) drives). NVM / storage device 1308 may include storage resources that are physically part of a device on which device 1300 is mounted, or that are accessible by that device but do not necessarily have to be part of that device. For example, NVM / storage device 1308 may be accessed via a network through (one or more) input / output devices 1310. One or more input / output devices 1310 may provide an interface for device 1300 to communicate with any other suitable device. Input / output devices 1310 may include communication components, pinyin components, sensor components, etc. Network interface 1312 may provide an interface for device 1300 to communicate via one or more networks. Device 1300 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G, 5G, etc., or combinations thereof.
[0254] In one embodiment, at least one of the processors 1302 may be logically packaged with one or more controllers (e.g., memory controller modules) of the control module 1304. In one embodiment, at least one of the processors 1302 may be logically packaged with one or more controllers of the control module 1304 to form a system-in-package (SiP). In one embodiment, at least one of the processors 1302 may be integrated with the logic of one or more controllers of the control module 1304 on the same die. In one embodiment, at least one of the processors 1302 may be integrated with the logic of one or more controllers of the control module 1304 on the same die to form a system-on-a-chip (SoC).
[0255] In various embodiments, device 1300 may be, but is not limited to, a server, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, device 1300 may have more or fewer components and / or different architectures. For example, in some embodiments, device 1300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0256] This application provides an electronic device, including: one or more processors; and one or more machine-readable media having instructions stored thereon, which, when executed by the one or more processors, cause the electronic device to perform one or more data processing methods as described in this application.
[0257] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0258] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0259] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0260] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0261] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0262] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.
[0263] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the element.
[0264] The data processing method and apparatus provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A data processing method, characterized in that, Applied to handheld terminals, including: At least one shooting mode is displayed on the screen of the handheld terminal; the shooting modes include at least: a shooting mode of a handheld terminal with a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal without a distance sensor mounted on a stabilizer, a shooting mode of a handheld terminal with a distance sensor, a shooting mode of a handheld terminal without a distance sensor, and a shooting mode of a handheld terminal without a distance sensor combined with a panoramic camera. Determine the selected target shooting mode from the displayed multiple shooting modes; The target shooting strategy adapted to the target shooting mode and the shooting device corresponding to the physical form of the target shooting mode are used to capture two-dimensional images of three-dimensional space and depth information of the two-dimensional images.
2. The method according to claim 1, characterized in that, Displaying at least one shooting mode on the screen of the handheld terminal includes: Determine whether the handheld terminal has a distance sensor and determine whether the handheld terminal is mounted on a stabilizer; When the handheld terminal has a distance sensor and is mounted on a stabilizer, at least the shooting mode of the handheld terminal with a distance sensor mounted on the stabilizer is displayed on the screen of the handheld terminal; or, When the handheld terminal has a distance sensor and is not mounted on a stabilizer, at least the shooting mode of the handheld terminal with the distance sensor is displayed on the screen of the handheld terminal; or, When the handheld terminal does not have a distance sensor and is mounted on a stabilizer, at least the shooting mode of the handheld terminal without a distance sensor mounted on the stabilizer is displayed on the screen of the handheld terminal; or, In the case where the handheld terminal does not have a distance sensor and is not mounted on a stabilizer, at least the shooting mode of the handheld terminal without a distance sensor and / or the shooting mode of the handheld terminal without a distance sensor combined with a panoramic camera are displayed on the screen of the handheld terminal.
3. The method according to claim 2, characterized in that, The provision of at least the shooting mode of a handheld terminal without a distance sensor and / or the shooting mode of a handheld terminal without a distance sensor combined with a panoramic camera displayed on the screen of the handheld terminal includes: Determine whether the handheld terminal is in communication connection with the panoramic camera; In the absence of a communication connection between the handheld terminal and the panoramic camera, at least the shooting mode of the handheld terminal without a distance sensor is displayed on the screen of the handheld terminal; or, When there is a communication connection between the handheld terminal and the panoramic camera, at least the shooting mode of the handheld terminal without a distance sensor and / or the shooting mode of the handheld terminal without a distance sensor combined with the panoramic camera are displayed on the screen of the handheld terminal.
4. The method according to claim 2 or 3, characterized in that, The provision of at least the shooting mode of a handheld terminal without a distance sensor and / or the shooting mode of a handheld terminal without a distance sensor combined with a panoramic camera displayed on the screen of the handheld terminal includes: The shooting mode of a handheld terminal without a distance sensor is displayed on the screen of the handheld terminal in a first display style, and the shooting mode of a handheld terminal without a distance sensor combined with a panoramic camera is displayed on the screen of the handheld terminal in a second display style. The salience of the first display style is less than that of the second display style.
5. The method according to claim 2, characterized in that, The steps of determining whether the handheld terminal has a distance sensor and determining whether the handheld terminal is mounted on a stabilizer include: The screen displays configuration options for the distance sensor and the stabilizer. When the distance sensor configuration option is selected and the stabilization configuration option is selected, it is determined that the handheld terminal has a distance sensor and is mounted on a stabilizer; or, If the distance sensor configuration option is selected and the stabilization configuration option is not selected, it is determined that the handheld terminal has a distance sensor but is not mounted on a stabilizer; or, If the distance sensor configuration option is not selected and the stabilization configuration option is selected, it is determined that the handheld terminal does not have a distance sensor and is mounted on a stabilizer; or, If the distance sensor configuration option is not selected and the stabilization configuration option is not selected, it is determined that the handheld terminal does not have a distance sensor and is not mounted on a stabilizer.
6. The method according to claim 2, characterized in that, Determining whether the handheld terminal is mounted on the stabilizer includes: Determine whether the handheld terminal is communicatively connected to the stabilizer; When the handheld terminal is communicatively connected to the stabilizer, it is determined that the handheld terminal is mounted on the stabilizer; or, If the handheld terminal and the stabilizer are not in a communication connection, it is determined that the handheld terminal is not mounted on the stabilizer.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the data processing method as described in any one of claims 1 to 6.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the data processing method as described in any one of claims 1 to 6.