Image projection method, device, system, electronic device and storage medium

By determining the corresponding points and depth values ​​of feature points in the 3D scene screenshot, directly calculating the projection matrix, and projecting the camera image into the 3D scene, the cumbersome operation problems in the existing technology are solved and a simpler and more accurate image projection effect is achieved.

CN117834824BActive Publication Date: 2025-09-16HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311869062.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-09-16
Estimated Expiration
2043-12-29

AI Technical Summary

Technical Problem

When projecting images taken by a camera into a three-dimensional scene, the existing technology requires manual calibration of the camera's working parameters and multiple changes of angles in the three-dimensional scene to determine the corresponding points of feature points. The operation steps are cumbersome and depend on the user's operating accuracy, resulting in a poor user experience.

Method used

By determining the corresponding points of the feature points of the target image in the screenshot of the 3D scene, combining the pixel coordinates and depth values ​​of the feature points, the projection matrix is ​​directly calculated and the image is projected into the 3D scene, avoiding manual calibration and multiple angle transformations.

Benefits of technology

The operation steps are simplified, the accuracy of feature point correspondence and the convenience of the projection process are improved, better image fusion effects are achieved, and the user experience is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117834824B_ABST
    Figure CN117834824B_ABST
Patent Text Reader

Abstract

The present application discloses an image projection method, device, system, electronic device and storage medium, which relates to the field of image processing technology and is used to reduce the user's operation steps in the process of projecting a target image into a three-dimensional scene, and quickly determine the corresponding points of feature points in the target image in the three-dimensional scene. The method includes: receiving a first calibration operation of a user on multiple feature points on a target image, determining the first pixel coordinates of each feature point; displaying a screenshot of the three-dimensional scene; determining the corresponding point of each feature point in the screenshot, and determining the second pixel coordinates of each corresponding point in the screenshot and the depth value of the corresponding point of the feature point in the three-dimensional scene; determining the projection matrix of the target camera based on the first pixel coordinate of each feature point, the second pixel coordinates of the corresponding point of the feature point, and the depth value; projecting the target image into the three-dimensional scene based on the projection matrix of the target camera, and displaying the projected three-dimensional scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image projection method, device, system, electronic device, and storage medium. Background Art

[0002] Videos captured by cameras installed at traffic intersections, key areas, and other locations play an irreplaceable and important role in daily management, emergency response, and event backtracking. However, the locations of individual cameras are often scattered, making it difficult to observe the videos captured by each camera from a holistic perspective.

[0003] At present, through 3D video fusion technology, the video images of each point can be seamlessly integrated with the 3D scene for splicing and display, and video retrieval and scheduling, video point switching, etc. can be quickly performed based on the 3D scene. This provides a more intuitive and convenient visual operation experience for indoor fixed-point or large-scale outdoor video detection or mobile position detection.

[0004] However, the three-dimensional video fusion technology requires that the video image captured by the camera be accurately projected to the corresponding position in the three-dimensional scene, which requires that the two-dimensional coordinates of the image captured by the camera be matched with the three-dimensional coordinates in the three-dimensional scene to determine the projection matrix of the camera. At present, there is a method of manually calibrating the corresponding points of the feature points on the image captured by the camera in the three-dimensional scene by determining the working parameters of the camera (for example, distortion parameters, coordinates, shooting posture, etc.), and then determining the projection matrix of the camera based on the two-dimensional coordinates of the feature points on the image captured by the camera and the three-dimensional coordinates of the corresponding points. This method not only requires obtaining the working parameters of the camera, but also requires the user to constantly change the angle in the three-dimensional scene to check whether the corresponding points at different angles completely overlap with the feature points, so as to more accurately determine the corresponding points of the feature points on the image captured by the camera. The user operation steps are relatively large, and the accuracy of the corresponding points of the feature points depends on the user repeatedly changing the angle of the three-dimensional scene to check whether the feature points and the corresponding points completely overlap, and the user's operation experience is not good. Summary of the Invention

[0005] The present application provides an image projection method, device, system, electronic device, and storage medium for reducing the user's operation steps in the process of projecting a target image into a three-dimensional scene, and quickly and accurately determining the corresponding points of feature points in the target image in the three-dimensional scene.

[0006] To achieve the above technical objectives, this application adopts the following technical solutions:

[0007] In a first aspect, an embodiment of the present application provides an image projection method, the method comprising:

[0008] Display the target image taken by the target camera;

[0009] receiving a first calibration operation of a user on a plurality of feature points on a target image, and determining a first pixel coordinate of each feature point;

[0010] Displaying a screenshot of the three-dimensional scene, where the area represented by the screenshot includes the area represented by the target image;

[0011] Determine the corresponding point of each feature point in the screenshot, and determine the second pixel coordinate of each corresponding point in the screenshot and the depth value of the corresponding point of the feature point in the three-dimensional scene;

[0012] Determine a projection matrix of the target camera based on the first pixel coordinates of each feature point, the second pixel coordinates of the corresponding point of the feature point, and the depth value;

[0013] Based on the projection matrix of the target camera, the target image is projected into the 3D scene and the projected 3D scene is displayed.

[0014] The technical solution provided by this application provides at least the following beneficial effects: During the process of projecting a target image into a three-dimensional scene, by determining the corresponding points of feature points in the target image from a screenshot of the three-dimensional scene, the user no longer needs to constantly change the angle in the three-dimensional scene to determine the corresponding points of the feature points in the target image, making the operation process simpler and more convenient, and determining the corresponding points of the feature points more accurate. When determining the corresponding points of the feature points, the depth value of the corresponding points in the three-dimensional scene is also determined. This allows the projection matrix of the target camera to be accurately determined by combining the first pixel coordinates of the feature points, the second pixel coordinates of the corresponding points of the feature points, and the depth value of the corresponding points, resulting in a better fusion effect when fusing the image captured by the target camera with the three-dimensional scene. This process does not require obtaining the operating parameters of the camera, nor does it require projecting the image captured by the camera to the corresponding position in the three-dimensional scene through texture projection based on the camera's operating parameters. Instead, this application directly uses the projection matrix and renders the image captured by the camera to the corresponding position in the three-dimensional scene through post-processing. This eliminates the need to determine the corresponding texture coordinates of the image captured by the camera in the three-dimensional scene, making the calculation process during the projection process more convenient and faster.

[0015] In one possible implementation, displaying a screenshot of a three-dimensional scene includes: displaying the three-dimensional scene in a three-dimensional scene display window; in response to receiving a screenshot operation, capturing an image of the three-dimensional scene displayed in the current three-dimensional scene display window to obtain a screenshot, and displaying the screenshot in the screenshot display window, the screenshot operation being performed when the area indicated by the three-dimensional scene displayed in the current three-dimensional scene display window includes the area indicated by the target image.

[0016] In one possible implementation, the method further includes: obtaining a depth map corresponding to the screenshot, the depth map being used to characterize the depth value of each pixel point in the screenshot; the depth value of the corresponding point of the feature point in the three-dimensional scene is determined in the following manner: determining the pixel point corresponding to the corresponding point of the feature point in the depth map, and determining the depth value of the pixel point as the depth value of the corresponding point of the feature point in the three-dimensional scene.

[0017] In one possible implementation, determining the corresponding point of each feature point in the screenshot includes: determining the corresponding point of each feature point in the screenshot in response to receiving a second calibration operation of the user on the corresponding point of the feature point in the screenshot; or, comparing the screenshot with the target image, identifying the corresponding point of each feature point from the screenshot.

[0018] In one possible implementation, the projection matrix of the target camera is determined based on the first pixel coordinates of each feature point, the second pixel coordinates of the corresponding point of the feature point, and the depth value, including: for each feature point, the NDC coordinates of the corresponding point are composed according to the second pixel coordinates of the corresponding point of the feature point, the depth value of the corresponding point of the feature point in the depth map, and a preset value, the NDC coordinates are used to represent the coordinates of the corresponding point in the converted coordinate system between the projection coordinate system and the two-dimensional coordinate system where the second pixel coordinates are located, the projection coordinate system is a three-dimensional coordinate system with the target camera as the origin and including the shooting range of the target camera; the NDC coordinates are converted into projection coordinates, the projection coordinates are used to represent the coordinates of the corresponding point in the projection coordinate system; the projection coordinates are converted into three-dimensional scene coordinates, the three-dimensional scene coordinates are the coordinates of the corresponding point in the spatial coordinate system of the three-dimensional scene; based on the first pixel coordinates of each feature point and the three-dimensional scene coordinates of the corresponding point of each feature point, the projection matrix of the target camera is determined.

[0019] In one possible implementation, the second pixel coordinates of the corresponding point of the feature point include pixel width and pixel height, and the preset value is 1; the NDC coordinates of the corresponding point are composed according to the second pixel coordinates of the corresponding point of the feature point, the depth value of the corresponding point of the feature point in the depth map, and the preset value, including: using the value obtained after performing a range processing operation on the pixel width as the first coordinate value of the NDC coordinate, using the value obtained after performing a range processing operation on the pixel height as the second coordinate value of the NDC coordinate, and using the value obtained after performing a range processing operation on the depth value as the third coordinate value of the NDC coordinate, the range processing operation is used to process the value into a value that meets the NDC coordinate range condition; based on the first coordinate value, the second coordinate value, the third coordinate value and 1, the NDC coordinates of the corresponding point are determined.

[0020] In a second aspect, the present application provides an image projection device, comprising:

[0021] A display module, used for displaying a target image captured by a target camera;

[0022] a processing module, configured to receive a first calibration operation of a user on a plurality of feature points on a target image, and determine a first pixel coordinate of each feature point;

[0023] The display module is further configured to display a screenshot of the three-dimensional scene, wherein the area represented by the screenshot includes the area represented by the target image;

[0024] The processing module is further configured to determine a corresponding point of each feature point in the screenshot, and determine a second pixel coordinate of each corresponding point in the screenshot and a depth value of the corresponding point of the feature point in the three-dimensional scene;

[0025] The processing module is further configured to determine a projection matrix of the target camera based on the first pixel coordinate of each feature point, the second pixel coordinate of a corresponding point of the feature point, and the depth value;

[0026] The processing module is further used to project the target image into a three-dimensional scene based on the projection matrix of the target camera, and display the projected three-dimensional scene.

[0027] In one possible implementation, the display module is specifically used to display a three-dimensional scene in a three-dimensional scene display window; the processing module is also used to, in response to receiving a screenshot operation, capture an image of the three-dimensional scene displayed in the current three-dimensional scene display window to obtain a screenshot; the display module is specifically used to display the screenshot in the screenshot display window, and the screenshot operation is performed when the area indicated by the three-dimensional scene displayed in the current three-dimensional scene display window includes the area indicated by the target image.

[0028] In one possible implementation, the processing module is also used to obtain a depth map corresponding to the screenshot, and the depth map is used to represent the depth value of each pixel point in the screenshot; the processing module is specifically used to determine the pixel point corresponding to the corresponding point of the feature point in the depth map, and determine the depth value of the pixel point as the depth value of the corresponding point of the feature point in the three-dimensional scene.

[0029] In one possible implementation, the processing module is specifically used to: determine the corresponding points of each feature point in the screenshot in response to receiving a second calibration operation of the user on the corresponding points of the feature points in the screenshot; or, compare the screenshot with the target image to identify the corresponding points of each feature point from the screenshot.

[0030] In one possible implementation, the processing module is specifically used to: for each feature point, form the NDC coordinates of the corresponding point based on the second pixel coordinates of the corresponding point of the feature point, the depth value of the corresponding point of the feature point in the depth map, and a preset value, the NDC coordinates are used to represent the coordinates of the corresponding point in the converted coordinate system between the projection coordinate system and the two-dimensional coordinate system where the second pixel coordinates are located, the projection coordinate system is a three-dimensional coordinate system with the target camera as the origin and including the shooting range of the target camera; convert the NDC coordinates into projection coordinates, the projection coordinates are used to represent the coordinates of the corresponding point in the projection coordinate system; convert the projection coordinates into three-dimensional scene coordinates, the three-dimensional scene coordinates are the coordinates of the corresponding point in the spatial coordinate system of the three-dimensional scene; determine the projection matrix of the target camera based on the first pixel coordinates of each feature point and the three-dimensional scene coordinates of the corresponding point of each feature point.

[0031] In one possible implementation, the second pixel coordinates of the corresponding point of the feature point include pixel width and pixel height, and the preset value is 1; the processing module is specifically used to: use the value obtained after performing a range processing operation on the pixel width as the first coordinate value of the NDC coordinate, use the value obtained after performing a range processing operation on the pixel height as the second coordinate value of the NDC coordinate, and use the value obtained after performing a range processing operation on the depth value as the third coordinate value of the NDC coordinate, and the range processing operation is used to process the value into a value that meets the NDC coordinate range condition; based on the first coordinate value, the second coordinate value, the third coordinate value and 1, determine the NDC coordinates of the corresponding point.

[0032] In a third aspect, the present application provides an electronic device comprising: one or more processors; one or more memories; wherein the one or more memories are used to store computer program code, the computer program code including computer instructions, and when the one or more processors execute the computer instructions, the electronic device performs any one of the image projection methods provided in the first aspect above.

[0033] In a fourth aspect, the present application provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are executed on a computer, the computer executes any one of the image projection methods provided in the first aspect.

[0034] In a fifth aspect, the present application provides a computer program product, which includes computer instructions. When the computer instructions are executed on a computer, the computer executes the image projection method as described in the first aspect and any possible design thereof.

[0035] For the specific descriptions of the second to fifth aspects and their various implementations in this application, reference can be made to the detailed descriptions in the first aspect and its various implementations; and for the beneficial effects of the second to fifth aspects and their various implementations, reference can be made to the analysis of the beneficial effects in the first aspect and its various implementations, which will not be repeated here.

[0036] These and other aspects of the present application will become more readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic structural diagram of an image projection system applicable to an image projection method provided in an embodiment of the present application;

[0038] Figure 2 A schematic diagram of the hardware composition of a computing device provided in an embodiment of the present application;

[0039] Figure 3 A flowchart of an image projection method provided in an embodiment of the present application;

[0040] Figure 4 Schematic diagram of an application scenario of an image projection method provided in an embodiment of the present application Figure 1 ;

[0041] Figure 5 Schematic diagram of an application scenario of an image projection method provided in an embodiment of the present application Figure 2 ;

[0042] Figure 6 Schematic diagram of an application scenario of an image projection method provided in an embodiment of the present application Figure 3 ;

[0043] Figure 7 Schematic diagram of an application scenario of an image projection method provided in an embodiment of the present application Figure 4 ;

[0044] Figure 8 Schematic diagram of an application scenario of an image projection method provided in an embodiment of the present application Figure 5 ;

[0045] Figure 9 Schematic diagram of an application scenario of an image projection method provided in an embodiment of the present application Figure 6 ;

[0046] Figure 10 Schematic diagram of an application scenario of an image projection method provided in an embodiment of the present application Figure 7 ;

[0047] Figure 11 Schematic diagram of an application scenario of an image projection method provided in an embodiment of the present application Figure 8;

[0048] Figure 12 Schematic diagram of an application scenario of an image projection method provided in an embodiment of the present application Figure 9 ;

[0049] Figure 13 A schematic structural diagram of an image projection device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0050] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0051] It should be noted that, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. To be precise, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete way. The terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise specified, "multiple" means two or more.

[0052] Videos captured by cameras installed at traffic intersections, key areas, and other locations play an irreplaceable and important role in daily management, emergency response, and event backtracking. However, the locations of individual cameras are often scattered, making it difficult to observe the videos captured by each camera from a holistic perspective.

[0053] At present, through 3D video fusion technology, the video images of each point can be seamlessly integrated with the 3D scene for splicing and display, and video retrieval and scheduling, video point switching, etc. can be quickly performed based on the 3D scene. This provides a more intuitive and convenient visual operation experience for indoor fixed-point or large-scale outdoor video detection or mobile position detection.

[0054] However, the three-dimensional video fusion technology requires that the video image captured by the camera be accurately projected to the corresponding position in the three-dimensional scene, which requires that the two-dimensional coordinates of the image captured by the camera be matched with the three-dimensional coordinates in the three-dimensional scene to determine the projection matrix of the camera. At present, there is a method of manually calibrating the corresponding points of the feature points on the image captured by the camera in the three-dimensional scene by determining the working parameters of the camera (for example, distortion parameters, coordinates, shooting posture, etc.), and then determining the projection matrix of the camera based on the two-dimensional coordinates of the feature points on the image captured by the camera and the three-dimensional coordinates of the corresponding points. This method not only requires obtaining the working parameters of the camera, but also requires the user to constantly change the angle in the three-dimensional scene to check whether the corresponding points at different angles completely overlap with the feature points, so as to more accurately determine the corresponding points of the feature points on the image captured by the camera. The user operation steps are relatively large, and the accuracy of the corresponding points of the feature points depends on the user repeatedly changing the angle of the three-dimensional scene to check whether the feature points and the corresponding points completely overlap, and the user's operation experience is not good.

[0055] In this regard, an embodiment of the present application provides an image projection method. In the process of projecting a target image into a three-dimensional scene, by determining the corresponding points of the feature points in the target image in a screenshot of the three-dimensional scene, the user does not need to constantly change the angle in the three-dimensional scene to determine the corresponding points of the feature points in the target image in the three-dimensional scene. The operation process is simpler and more convenient, and the determination of the corresponding points of the feature points is also more accurate. When determining the corresponding points of the feature points, the depth value of the corresponding points in the three-dimensional scene is also determined. Therefore, the projection matrix of the target camera can be accurately determined by combining the first pixel coordinates of the feature points, the second pixel coordinates of the corresponding points of the feature points, and the depth value of the corresponding points. When the image captured by the target camera is fused with the three-dimensional scene, a better fusion effect can be obtained. This process does not require obtaining the operating parameters of the camera, and does not require projecting the image captured by the camera to the corresponding position in the three-dimensional scene through texture projection based on the operating parameters of the camera. The present application directly uses the projection matrix and renders the image captured by the camera to the corresponding position in the three-dimensional scene through post-processing. There is no need to determine the corresponding texture coordinates of the image captured by the camera in the three-dimensional scene. The calculation process during the projection process is also more convenient and quick.

[0056] Please refer to Figure 1 , which shows the image projection system to which the image projection method provided by this application is applicable. Figure 1 As shown, the image projection system 1 includes: a plurality of cameras 10 and an electronic device 20 .

[0057] A communication connection is established between the camera 10 and the electronic device 20. It should be understood that the connection can be a wireless connection, such as a Bluetooth connection or a wireless fidelity (Wi-Fi) connection, or a wired connection, such as an optical fiber connection, without limitation. For example, the camera 10 and the electronic device 20 can be connected to the Internet via a router, thereby enabling communication between the camera 10 and the electronic device 20.

[0058] In some embodiments, the camera 10 is used to capture video images and send the video images to the electronic device 20 when the video images need to be projected into a three-dimensional scene.

[0059] In some embodiments, the electronic device 20 stores a three-dimensional scene of a specific area, and the three-dimensional scene may be obtained by pre-modeling the specific area.

[0060] In some embodiments, the electronic device 20 further stores a depth value of each point in the three-dimensional scene when modeling a specific area.

[0061] In some embodiments, the electronic device 20 is configured to project the video image captured by the camera 10 into a three-dimensional scene.

[0062] Specifically, the electronic device 20 displays the target image captured by the target camera; receives the user's first calibration operation on multiple feature points on the target image, and determines the first pixel coordinates of each feature point; displays a screenshot of the three-dimensional scene, and the area represented by the screenshot includes the area represented by the target image; determines the corresponding point of each feature point in the screenshot, and determines the second pixel coordinates of each corresponding point in the screenshot and the depth value of the corresponding point of the feature point in the three-dimensional scene; determines the projection matrix of the target camera based on the first pixel coordinate of each feature point, the second pixel coordinates of the corresponding point of the feature point, and the depth value; finally, the electronic device 20 projects the target image into the three-dimensional scene based on the projection matrix of the target camera, and displays the projected three-dimensional scene.

[0063] In some embodiments, when the electronic device 20 displays a three-dimensional scene, it renders the displayed image in real time. The rendering process can obtain a depth map corresponding to the displayed image. Therefore, when a screenshot of the three-dimensional scene is obtained, a depth map corresponding to the screenshot can be obtained. In this way, the electronic device can determine the depth value of the corresponding point based on the pixel point corresponding to the corresponding point marked by the user in the screenshot in the depth map.

[0064] In some embodiments, the electronic device 20 may project the video images captured by each camera 10 into a three-dimensional scene.

[0065] In some embodiments, the electronic device 20 may include a human-computer interaction device for displaying the image projection results and the projected three-dimensional scene to the user, and for receiving user calibration operations for feature points and / or corresponding points of feature points. The human-computer interaction device may be a display, such as a liquid crystal display or an organic light-emitting diode (OLED) display. The specific type, size, and resolution of the display are not limited.

[0066] In some embodiments, the electronic device 20 may be a single server or a server cluster, or the electronic device 20 may be a terminal device, such as a personal computer (PC), a notebook computer, a mobile device, a tablet computer, a laptop computer, etc. The embodiments of the present application do not limit the specific form of the electronic device 20.

[0067] In some embodiments, the image projection system 1 may also include multiple electronic devices 20 to facilitate multiple users to project images to multiple locations at the same time.

[0068] The hardware structure of the electronic device 20 includes Figure 2 The computing device shown in FIG. Figure 2 Taking the computing device shown as an example, the hardware structure of the electronic device 20 is introduced.

[0069] like Figure 2 As shown, the computing device may include a processor 301 , a memory 302 , a communication interface 303 , and a bus 304 . The processor 301 , the memory 302 , and the communication interface 303 may be connected via the bus 304 .

[0070] Processor 301 is the control center of the computing device and can be a single processor or a collective term for multiple processing elements. For example, processor 301 can be a general-purpose central processing unit (CPU) or other general-purpose processor. A general-purpose processor can be a microprocessor or any conventional processor.

[0071] As an embodiment, the processor 301 may include one or more CPUs, such as Figure 3 CPU 0 and CPU 1 are shown in Figure 1.

[0072] The memory 302 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0073] In one possible implementation, memory 302 may exist independently of processor 301 and may be connected to processor 301 via bus 304 for storing instructions or program code. When processor 301 calls and executes the instructions or program code stored in memory 302, the model deployment method provided in the embodiments of the present application can be implemented.

[0074] In another possible implementation, the memory 302 may also be integrated with the processor 301 .

[0075] The communication interface 303 is used to connect the computing device to other devices via a communication network, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 303 may include a receiving unit for receiving data and a sending unit for sending data.

[0076] The bus 304 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 2 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0077] It should be pointed out that Figure 2 The structure shown in the figure does not constitute a limitation on the computing device, except Figure 2In addition to the components shown, the computing device may include more or fewer components than shown, or combine certain components, or arrange the components differently.

[0078] The implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0079] like Figure 3 As shown, an embodiment of the present application provides an image projection method, which can be performed by the above-mentioned electronic device 20, and the method includes the following steps:

[0080] S101: Display a target image captured by a target camera.

[0081] In some embodiments, the electronic device displays an identification information input control; then receives identification information of the target camera input by the user based on the identification information input control, obtains the target image taken by the target camera, and displays the target image taken by the target camera in the image display window.

[0082] Optionally, the electronic device obtains a video stream captured by a target camera, and the user specifies one frame of image as the target image, or the electronic device randomly selects one frame of image from the video stream captured by the target camera as the target image.

[0083] For example, the identification information of the target camera may be the number, name, identity document (ID) number of the target camera, the uniform resource locator (URL) of the target camera, or the Internet protocol (IP) address of the target camera. Figure 4 As shown, an identification information input control is displayed in the video fusion editor of the electronic device. The user enters "cube3.mp4" in it, and the electronic device obtains and displays a frame of image taken by a camera with the identification information "cube3.mp4".

[0084] In this way, the user can specify the target camera for image projection to obtain an image projection result that meets the user's needs.

[0085] S102: Receive a first calibration operation of a user on a plurality of feature points on a target image, and determine a first pixel coordinate of each feature point.

[0086] The feature points should be easily identifiable pixels, which facilitates the subsequent selection of corresponding points. To ensure the accuracy of image projection, it is usually necessary to select more than 6 non-coplanar feature points in the target image, and the number of feature points selected is generally 6-10.

[0087] For example, Figure 5 As shown in , the target image captured by the target camera is a display image of a cube, so the user can select the six vertices of the cube in the target image as feature points, such as Figure 5 The “1, 2, 3, 4, 5, 6” marked in the figure are the labels of the feature points.

[0088] In some embodiments, after the electronic device determines the first pixel coordinates of each feature point, it normalizes the first pixel of each feature point. For example, if the pixel size of the target image is 1920×1080, and the pixel coordinates of a feature point in the target image are (192, 108), then the pixel coordinates are normalized to obtain (1 1 0, 1 1 0). In this way, normalizing the pixel coordinates of the feature points can facilitate the subsequent calculation of the projection matrix of the target camera based on the first pixel coordinates.

[0089] For example, Figure 5 As shown, after the user marks the feature point, the electronic device determines the first pixel coordinates of the feature point, normalizes the first pixel coordinates, and displays the normalized first pixel coordinates.

[0090] S103: Display a screenshot of the three-dimensional scene.

[0091] The area represented by the screenshot includes the area represented by the target image.

[0092] In some embodiments, the screenshot has a corresponding depth map, which is used to represent the depth value of each pixel point in the screenshot, so that the electronic device can determine the depth value of each corresponding point based on the depth map corresponding to the screenshot in the subsequent process.

[0093] For example, a screenshot of a 3D scene can be as follows: Figure 6 As shown in (a), the corresponding depth map can be Figure 6 As shown in (b) in .

[0094] In practical applications, the depth value can be recorded in the R channel, G channel or B channel of the pixel. When the depth value is recorded in the R channel of the pixel, Figure 6 The observed color of the cube in (b) is red.

[0095] In some embodiments, the electronic device displays a 3D scene in a 3D scene display window; in response to receiving a screenshot operation, the electronic device captures an image of the 3D scene displayed in the current 3D scene display window to obtain a screenshot, and displays the screenshot in the screenshot display window. The screenshot operation is performed when the area indicated by the 3D scene displayed in the current 3D scene display window includes the area indicated by the target image.

[0096] Optionally, the 3D scene display window and the screenshot display window may be the same display window, or the 3D scene display window and the screenshot display window may be different display windows.

[0097] For example, Figure 7 As shown, a three-dimensional scene is displayed in the three-dimensional scene display window of the electronic device, and the user controls the electronic device to change the display angle of the three-dimensional scene. When it is determined that the video image taken by the target camera is to be projected in an area represented by a certain three-dimensional scene, the "Create Projection Point" control in the electronic device is clicked to trigger the electronic device to capture the image of the three-dimensional scene displayed in the current three-dimensional scene display window. At this time, the screenshot operation is the operation of the user clicking the "Create Projection Point" control in the electronic device.

[0098] As another example, when the user needs to mark the corresponding points of the feature points in the screenshot of the three-dimensional scene, the user can click the "scene mark" control in the electronic device, such as Figure 8 As shown, the electronic device displays a screenshot of the three-dimensional scene in the screenshot display window, and then the user clicks the "scene marking" control to mark the corresponding points of each feature point in the screenshot.

[0099] In some embodiments, after the electronic device determines the second pixel coordinates of the corresponding point, it normalizes the second pixel coordinates to facilitate subsequent determination of the three-dimensional coordinates of the corresponding point based on the normalized second pixel coordinates. Figure 8 As shown, after the user selects the corresponding point, the electronic device displays the second pixel coordinates of the corresponding point after normalization.

[0100] Optionally, the screenshot operation is issued by the user, and the screenshot area indicated by the screenshot operation includes an area represented by the target image.

[0101] For example, when the user clicks the "Create Projection Point" control in the electronic device, Figure 9 As shown, the electronic device displays an area selection control, and the user can control the area selection control to determine the area that needs to be screenshotted, and take a screenshot of the area selected by the user. At this time, the screenshot operation is that the user clicks the "Create Projection Point" control in the electronic device, and controls the area selection control to determine the area that needs to be screenshotted.

[0102] It should be understood that when projecting the target image into a three-dimensional scene, a screenshot of the three-dimensional scene is only needed to mark the corresponding points of the feature points. The final projection display effect of the target image is unrelated to the screenshot of the three-dimensional scene. In other words, in actual applications, it is not necessary to pay attention to whether the screenshot of the three-dimensional scene is a full or partial screenshot of the three-dimensional scene display window. It is only necessary that the area represented by the screenshot include the area represented by the target image. If a partial screenshot method is used, the user needs to manually select the screenshot area, which increases the user operation steps and has no effect on the final display effect. Therefore, in actual applications, the method of taking a screenshot of the entire three-dimensional scene display window is usually adopted to determine the screenshot of the three-dimensional scene.

[0103] In this way, if it is necessary to project the image taken by the target camera in a three-dimensional scene, by taking a screenshot of the three-dimensional scene displayed in the current three-dimensional scene display window, a screenshot of the three-dimensional scene containing the area indicated by the target image can be taken. The user's subsequent process of determining the corresponding points of each feature point in the screenshot is also simpler, and there is no need to transform the three-dimensional scene from multiple angles to determine the corresponding points of each feature point. The corresponding points finally determined are not only more accurate, but also reduce the user's operation steps. When the target image is subsequently projected into the three-dimensional scene, better projection results can be obtained, which also improves the user experience.

[0104] S104 , determining a corresponding point of each feature point in the screenshot, and determining a second pixel coordinate of each corresponding point in the screenshot and a depth value of the corresponding point of the feature point in the three-dimensional scene.

[0105] In some embodiments, the electronic device determines the corresponding points of each feature point in the screenshot in response to receiving a second calibration operation by the user on the corresponding points of the feature points in the screenshot; or, the electronic device compares the screenshot with the target image and identifies the corresponding points of each feature point from the screenshot.

[0106] For example, Figure 8 As shown, the user marks the vertices of the cube in the screenshot display window as corresponding points of the feature points. Alternatively, because feature points are easy to identify, and the feature points in the area represented by the target image should correspond to the corresponding points in the area represented by the screenshot of the 3D scene, the user can pre-train a recognition algorithm that can identify feature points, and then identify the corresponding points of the feature points from the screenshot.

[0107] In this way, compared to constantly changing the angle to find the corresponding points of the feature points in the three-dimensional scene, the screenshot of the three-dimensional scene is a two-dimensional image. The process of users marking the corresponding points is more convenient and faster, and the marked corresponding points are more accurate, which can reduce the user's operation steps, save the user's operation time, and improve the user experience. At the same time, in the subsequent process, only the first pixel coordinates and the second pixel coordinates need to be calculated through the preset calculation logic to determine the three-dimensional coordinates of the corresponding points in the three-dimensional scene, and then determine the projection matrix of the target camera.

[0108] In some embodiments, the electronic device obtains a depth map corresponding to the screenshot, and then determines the pixel points corresponding to the corresponding points of the feature points in the depth map, and determines the depth value of the pixel points as the depth value of the corresponding points of the feature points in the three-dimensional scene.

[0109] It should be understood that the screenshot and the depth map are matched, and the pixels in the screenshot correspond one-to-one to the pixels in the depth map. After the user marks the corresponding point in the screenshot, the electronic device determines the pixel point corresponding to this corresponding point in the depth map, and obtains the depth map of this corresponding point.

[0110] In this way, since the electronic device can render the depth map of the electronic device display screen in real time when displaying the three-dimensional scene, the electronic device can quickly obtain the depth map corresponding to the screenshot of the three-dimensional scene, and then the electronic device can automatically calculate the three-dimensional scene coordinates of the corresponding point based on the depth value of the corresponding point, the pixel coordinates in the screenshot and other information, which facilitates the subsequent calculation of the projection matrix, so that the camera image can be projected into the three-dimensional scene based on the projection matrix.

[0111] In some embodiments, the electronic device determines the depth value of the corresponding point of the feature point from the depth value of each point in the three-dimensional scene stored when the three-dimensional scene is modeled.

[0112] For example, when a user marks a corresponding point in a screenshot, the electronic device identifies the point corresponding to the corresponding point in the three-dimensional scene from the pre-stored depth value of each point in the three-dimensional scene, and then finds the depth value of the corresponding point.

[0113] S105 : Determine a projection matrix of the target camera based on the first pixel coordinates of each feature point, the second pixel coordinates of the corresponding point of the feature point, and the depth value.

[0114] In some embodiments, for each feature point, the electronic device composes the NDC coordinates of the corresponding point based on the second pixel coordinates of the corresponding point of the feature point, the depth value of the corresponding point of the feature point in the depth map, and a preset value. The NDC coordinates are used to represent the coordinates of the corresponding point in the converted coordinate system between the projection coordinate system and the two-dimensional coordinate system where the second pixel coordinates are located. The projection coordinate system is a three-dimensional coordinate system with the target camera as the origin and including the shooting range of the target camera; the electronic device converts the NDC coordinates into projection coordinates, and the projection coordinates are used to represent the coordinates of the corresponding point in the projection coordinate system; the electronic device converts the projection coordinates into three-dimensional scene coordinates, and the three-dimensional scene coordinates are the coordinates of the corresponding point in the spatial coordinate system in the three-dimensional scene; the electronic device determines the projection matrix of the target camera based on the first pixel coordinates of each feature point and the three-dimensional scene coordinates of the corresponding point of each feature point.

[0115] The preset value is used to represent the dimension corresponding to the range of each coordinate value when coordinate conversion is performed using NDC coordinates. Usually, the preset value is 1.

[0116] In some embodiments, after obtaining the projection coordinates, the electronic device further performs homogenization processing on the projection coordinates to obtain homogenized projection coordinates, and converts the homogenized projection coordinates into three-dimensional scene coordinates.

[0117] Specifically, the electronic device uses the value obtained after performing a range processing operation on the pixel width as the first coordinate value of the NDC coordinate, uses the value obtained after performing a range processing operation on the pixel height as the second coordinate value of the NDC coordinate, and uses the value obtained after performing a range processing operation on the depth value as the third coordinate value of the NDC coordinate. The range processing operation is used to process the value into a value that meets the NDC coordinate range condition; based on the first coordinate value, the second coordinate value, the third coordinate value and 1, the NDC coordinate of the corresponding point is determined.

[0118] For example, U represents the normalized value of the pixel width in the second pixel coordinate of the corresponding point of the feature point, V represents the normalized value of the pixel height in the second pixel coordinate of the corresponding point of the feature point, and D represents the depth value of the corresponding point of the feature point in the depth map. The electronic device may perform range processing on the pixel width as follows: the electronic device normalizes U, and the obtained value has a value range of (0,1). Since the coordinate value in the ndc coordinate has a value range of (-1,1), the electronic device converts U into a value in the range of (-1,1) to obtain 2U-1. The same operation is performed on V and D to obtain 2V-1, 2(1-D)-1. That is, the ndc coordinates of the corresponding point can be expressed as:

[0119] ndc=(2U-1, 2V-1, 2(1-D)-1, 1)

[0120] For example, the projection matrix can be used to convert the ndc coordinates of the corresponding points into projection coordinates, that is, the matrix that converts the projection (View) coordinates in the three-dimensional scene into ndc coordinates, and then the projection coordinates are processed based on the preset homogenization value View.w to obtain the homogenized projection coordinates. The electronic device then uses the WorldToCameraMatrix, a matrix that converts 3D scene (World) coordinates to View coordinates, to convert the homogenized projection coordinates into the 3D scene coordinates of the corresponding point in the 3D scene. Finally, the electronic device uses the NormalDLT algorithm to calculate the World coordinates and the first pixel coordinates to obtain the projection matrix of the target camera.

[0121] In this way, through simple calculation logic, the first pixel coordinates and the second pixel coordinates can be calculated to determine the three-dimensional scene coordinates of the corresponding points in the three-dimensional scene, and then the projection matrix of the target camera can be determined. This process can be calculated by the electronic device itself, and the user does not need to manually determine the three-dimensional coordinates of the corresponding points of the feature points in the three-dimensional scene, which can improve the user experience.

[0122] S106 : Based on the projection matrix of the target camera, project the target image into the three-dimensional scene, and display the projected three-dimensional scene.

[0123] For example, Figure 10 As shown in the figure, the user can click the click management menu control in the electronic device, and a pop-up menu will display the visual fusion point ID, the identification information of the target camera, the coordinates of the target camera and the shooting angle. The user can also click the preview control to preview the fusion result of the target image and the 3D scene. Finally, after the target image is projected into the 3D scene, the projection effect can be as follows: Figure 11 As shown, the area with thinner lines is the shooting picture of the target camera, and the area with thicker lines is the three-dimensional scene.

[0124] In some embodiments, the electronic device obtains a video stream captured by a target camera; based on a projection matrix, projects each video image in the video stream into a three-dimensional scene, and displays the video stream projected in the three-dimensional scene.

[0125] In this way, the images captured by the target camera can be displayed in real time in the three-dimensional scene, making it easier for users to understand the area represented by the three-dimensional scene from a holistic perspective by combining the camera projection results at other points.

[0126] In some embodiments, the electronic device receives a user's drawing operation on a fusion area in a target image, determines the fusion area in the target image, and then, in response to receiving the user's fusion operation, projects the fusion area in the target image into a three-dimensional scene based on the projection matrix of the target camera.

[0127] For example, Figure 12 As shown, the user can mark multiple pixels in the target image in a clockwise order, and the electronic device connects the marked pixels in turn, and finally connects the formed area ( Figure 12 The area (defined by the dashed line) is the fusion area. The electronic device then receives a click on the "Video Fusion" control and projects the selected fusion area into the 3D scene. Alternatively, the user can mark multiple pixels in the target image in a counterclockwise order. The electronic device then sequentially connects the marked pixels to determine the fusion area.

[0128] In this way, the user can select the image in the target image that needs to be projected into the three-dimensional scene, which can better meet the user's usage needs.

[0129] Figure 3 The technical solution provided herein provides at least the following beneficial effects: During the process of projecting a target image into a three-dimensional scene, by determining the corresponding points of feature points in the target image in a screenshot of the three-dimensional scene, the user does not need to constantly change the angle in the three-dimensional scene to determine the corresponding points of the feature points in the target image in the three-dimensional scene, making the operation process simpler and more convenient, and determining the corresponding points of the feature points more accurate. When determining the corresponding points of the feature points, the depth value of the corresponding points in the three-dimensional scene is also determined. This allows the projection matrix of the target camera to be accurately determined by combining the first pixel coordinates of the feature points, the second pixel coordinates of the corresponding points of the feature points, and the depth value of the corresponding points, thereby achieving a better fusion effect when fusing the image captured by the target camera with the three-dimensional scene. This process does not require obtaining the operating parameters of the camera, nor does it require projecting the image captured by the camera to the corresponding position in the three-dimensional scene through texture projection based on the camera operating parameters. Instead, the present application directly uses the projection matrix and renders the image captured by the camera to the corresponding position in the three-dimensional scene through post-processing. This eliminates the need to determine the corresponding texture coordinates of the image captured by the camera in the three-dimensional scene, and the calculation process during the projection process is also more convenient and faster.

[0130] The following describes the actual application process of the image projection method provided by this solution from the perspective of the overall process:

[0131] First, if Figure 4As shown, the user clicks on the point creation control in the electronic device to create a new video fusion point, enters the identification information corresponding to the video fusion point (for example, the user enters the number and ID number of the video fusion point in the electronic device), and the stream address of the video stream to be fused (for example, the identification information of the target camera), and then clicks on the save control in the electronic device to complete the preparations for creating the video fusion point.

[0132] Afterwards, if Figure 5 As shown, the user clicks the video dot control in the electronic device, and a frame image of the previously taken video stream to be fused is displayed in the target image display window. The user selects more than 6 feature points in the image displayed in the target image display window.

[0133] like Figure 8 As shown, the user clicks the scene dot control in the electronic device, and a screenshot of the three-dimensional scene is displayed in the screenshot display window. The user selects the corresponding point of each feature point on the screenshot, and the electronic device automatically calculates the three-dimensional coordinates of each corresponding point.

[0134] The user clicks the save point control in the electronic device, and the electronic device automatically calculates the projection matrix of the target camera of the video fusion point.

[0135] The user clicks the mask dot control on the electronic device to draw the mask points in a clockwise or counterclockwise direction, thereby determining the fusion area.

[0136] like Figure 12 As shown, the user clicks on the control to generate the mask image in the electronic device and then clicks on the preview effect of showing or hiding the mask.

[0137] The user clicks the Save Point control to save the mask image.

[0138] The user clicks the Save All control to save the data determined by the above process.

[0139] like Figure 10 As shown, the user clicks the management menu control, and a pop-up menu displays the visual fusion point ID, the stream address of the video stream to be fused, the coordinates and shooting angle of the target camera. The user can also click the preview control to preview the fusion result of the video stream to be fused and the 3D scene.

[0140] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. It should be easy to realize that the technical goals in this field are combined with the units and algorithm steps of each example described in the embodiments disclosed herein, and the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technical goals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0141] like Figure 13 As shown, the embodiment of the present application further provides an image projection device, which is used in the image projection method shown in the above method embodiment. The image projection device 400 includes: a display module 401 and a processing module 402.

[0142] Among them, the display module 401 is used to display the target image captured by the target camera; the processing module 402 is used to receive the user's first calibration operation on multiple feature points on the target image, and determine the first pixel coordinates of each feature point; the display module 401 is also used to display a screenshot of the three-dimensional scene, and the area represented by the screenshot includes the area represented by the target image; the processing module 402 is also used to determine the corresponding point of each feature point in the screenshot, and determine the second pixel coordinates of each corresponding point in the screenshot and the depth value of the corresponding point of the feature point in the three-dimensional scene; the processing module 402 is also used to determine the projection matrix of the target camera based on the first pixel coordinate of each feature point, the second pixel coordinates of the corresponding point of the feature point and the depth value; the processing module 402 is also used to project the target image into the three-dimensional scene based on the projection matrix of the target camera, and display the projected three-dimensional scene.

[0143] In one possible implementation, the display module 401 is specifically used to display a three-dimensional scene in a three-dimensional scene display window; the processing module 402 is also used to, in response to receiving a screenshot operation, capture an image of the three-dimensional scene displayed in the current three-dimensional scene display window to obtain a screenshot; the display module 401 is specifically used to display the screenshot in the screenshot display window, and the screenshot operation is performed when the area indicated by the three-dimensional scene displayed in the current three-dimensional scene display window includes the area indicated by the target image.

[0144] In one possible implementation, the processing module 402 is also used to obtain a depth map corresponding to the screenshot, where the depth map is used to represent the depth value of each pixel in the screenshot; the processing module 402 is specifically used to determine the pixel points corresponding to the corresponding points of the feature points in the depth map, and determine the depth value of the pixel point as the depth value of the corresponding point of the feature point in the three-dimensional scene.

[0145] In one possible implementation, the processing module 402 is specifically used to: determine the corresponding points of each feature point in the screenshot in response to receiving a second calibration operation of the user on the corresponding points of the feature points in the screenshot; or, compare the screenshot with the target image to identify the corresponding points of each feature point from the screenshot.

[0146] In one possible implementation, the processing module 402 is specifically used to: for each feature point, form the NDC coordinates of the corresponding point based on the second pixel coordinates of the corresponding point of the feature point, the depth value of the corresponding point of the feature point in the depth map, and a preset value, the NDC coordinates are used to represent the coordinates of the corresponding point in the converted coordinate system between the projection coordinate system and the two-dimensional coordinate system where the second pixel coordinates are located, and the projection coordinate system is a three-dimensional coordinate system with the target camera as the origin and including the shooting range of the target camera; convert the NDC coordinates into projection coordinates, and the projection coordinates are used to represent the coordinates of the corresponding point in the projection coordinate system; convert the projection coordinates into three-dimensional scene coordinates, and the three-dimensional scene coordinates are the coordinates of the corresponding point in the spatial coordinate system of the three-dimensional scene; determine the projection matrix of the target camera based on the first pixel coordinates of each feature point and the three-dimensional scene coordinates of the corresponding point of each feature point.

[0147] In one possible implementation, the second pixel coordinates of the corresponding point of the feature point include pixel width and pixel height, and the preset value is 1; the processing module 402 is specifically used to: use the value obtained after performing a range processing operation on the pixel width as the first coordinate value of the NDC coordinate, use the value obtained after performing a range processing operation on the pixel height as the second coordinate value of the NDC coordinate, and use the value obtained after performing a range processing operation on the depth value as the third coordinate value of the NDC coordinate, and the range processing operation is used to process the value into a value that meets the NDC coordinate range condition; based on the first coordinate value, the second coordinate value, the third coordinate value and 1, determine the NDC coordinates of the corresponding point.

[0148] It should be noted that Figure 13 The module division described is illustrative and represents only one logical functional division. Actual implementations may employ different divisions. For example, two or more functions may be integrated into a single processing module. These integrated modules may be implemented as either hardware or software functional modules.

[0149] Another embodiment of the present application provides an electronic device comprising a memory and a processor; the memory and the processor are coupled; the memory is configured to store computer program code, the computer program code comprising computer instructions. When the processor executes the computer instructions, the electronic device executes each step of the method flow shown in the above method embodiment.

[0150] In actual implementation, the display module 401 and the processing module 402 can be implemented by the processor of the electronic device calling the computer program code in the memory. The specific execution process can be referred to the description of the image projection method above, which will not be repeated here.

[0151] Another embodiment of the present application further provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed on a computer, the computer executes each step performed by the electronic device in the method flow shown in the above method embodiment.

[0152] In another embodiment of the present application, a computer program product is provided. The computer program product includes computer instructions. When the computer instructions are executed on a computer, the steps performed by the electronic device in the method flow shown in the above method embodiment are executed at three levels.

[0153] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer execution instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more servers that can be integrated with the medium. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), etc.

[0154] The above is only a specific embodiment of the present application. Those skilled in the art may conceive of changes or substitutions based on the specific embodiment provided in this application, and all such changes or substitutions shall fall within the scope of protection of this application.

Claims

1. An image projection method, characterized in that: include: Display the target image taken by the target camera; receiving a first calibration operation of a user on a plurality of feature points on the target image, and determining a first pixel coordinate of each of the feature points; Displaying a screenshot of the three-dimensional scene, wherein the area represented by the screenshot includes the area represented by the target image; Determining a corresponding point of each of the feature points in the screenshot, and determining a second pixel coordinate of each of the corresponding points in the screenshot and a depth value of the corresponding point of the feature point in the three-dimensional scene; Determining a projection matrix of the target camera based on the first pixel coordinates of each of the feature points, the second pixel coordinates of a corresponding point of the feature point, and the depth value; Based on the projection matrix of the target camera, the target image is projected into the three-dimensional scene, and the projected three-dimensional scene is displayed.

2. The method according to claim 1, characterized in that The screenshot showing the three-dimensional scene includes: Displaying the three-dimensional scene in a three-dimensional scene display window; In response to receiving a screenshot operation, an image of the three-dimensional scene displayed in the current three-dimensional scene display window is captured to obtain the screenshot, and the screenshot is displayed in the screenshot display window. The screenshot operation is performed when the area indicated by the three-dimensional scene displayed in the current three-dimensional scene display window includes the area indicated by the target image.

3. The method according to claim 1, characterized in that The method further comprises: Obtaining a depth map corresponding to the screenshot, wherein the depth map is used to represent the depth value of each pixel in the screenshot; The depth value of the corresponding point of the feature point in the three-dimensional scene is determined by: A pixel point corresponding to the corresponding point of the feature point in the depth map is determined, and a depth value of the pixel point is determined as a depth value of the corresponding point of the feature point in the three-dimensional scene.

4. The method according to claim 1, wherein Determining the corresponding point of each feature point in the screenshot includes: In response to receiving a second calibration operation of the user on the corresponding point of the feature point in the screenshot, determining the corresponding point of each feature point in the screenshot; or, The screenshot is compared with the target image, and corresponding points of each feature point are identified from the screenshot.

5. The method according to any one of claims 1 to 4, characterized in that The determining the projection matrix of the target camera based on the first pixel coordinates of each feature point, the second pixel coordinates of the corresponding point of the feature point, and the depth value includes: For each of the feature points, composing an NDC coordinate of the corresponding point based on the second pixel coordinate of the corresponding point of the feature point, the depth value of the corresponding point of the feature point in the depth map, and a preset value, wherein the NDC coordinate is used to represent the coordinates of the corresponding point in a transformed coordinate system between a projected coordinate system and a two-dimensional coordinate system in which the second pixel coordinates are located, wherein the projected coordinate system is a three-dimensional coordinate system with the target camera as its origin and including a shooting range of the target camera; Converting the NDC coordinates into projection coordinates, where the projection coordinates are used to represent the coordinates of the corresponding points in the projection coordinate system; Converting the projection coordinates into three-dimensional scene coordinates, where the three-dimensional scene coordinates are coordinates of the corresponding point in a spatial coordinate system in the three-dimensional scene; The projection matrix of the target camera is determined based on the first pixel coordinates of each of the feature points and the three-dimensional scene coordinates of a corresponding point of each of the feature points.

6. An image projection device, characterized in that: include: A display module, used for displaying a target image captured by a target camera; a processing module, configured to receive a first calibration operation of a user on a plurality of feature points on the target image, and determine a first pixel coordinate of each of the feature points; The display module is further configured to display a screenshot of the three-dimensional scene, wherein the area represented by the screenshot includes the area represented by the target image; The processing module is further configured to determine a corresponding point of each of the feature points in the screenshot, and determine a second pixel coordinate of each corresponding point in the screenshot and a depth value of the corresponding point of the feature point in the three-dimensional scene; The processing module is further configured to determine a projection matrix of the target camera based on the first pixel coordinates of each feature point, the second pixel coordinates of a corresponding point of the feature point, and the depth value; The processing module is further configured to project the target image into the three-dimensional scene based on the projection matrix of the target camera, and display the projected three-dimensional scene.

7. An electronic device, characterized in that: include: one or more processors; one or more memories; The one or more memories are used to store computer program codes, and the computer program codes include computer instructions. When the one or more processors execute the computer instructions, the electronic device executes the image projection method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed on a computer, the computer is enabled to execute the image projection method according to any one of claims 1 to 5.

9. A computer program product, characterized in that The computer program product comprises computer instructions, and when the computer instructions are run on a computer, the computer is caused to perform the image projection method according to any one of claims 1 to 5.

10. An image projection system, characterized in that: The device comprises the electronic device according to claim 7, and at least one camera, wherein the camera is used to capture images, and the electronic device is used to project the images captured by the camera into a three-dimensional scene.

Citation Information

Patent Citations

  • Image data display method and device, electronic equipment and storage medium

    CN111325824A

  • Image processing method and device, electronic equipment and computer readable storage medium

    CN113706609A