Image acquisition equipment calibration method, electronic equipment, storage medium and product
By automatically controlling the movement of virtual image acquisition devices and matching images in the digital twin world, the problem of low camera calibration efficiency is solved and an efficient automated calibration process is achieved.
Patent Information
- Application Number
- CN202410316316.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2025-09-19
AI Technical Summary
In the existing technology, the camera calibration process of the digital twin world is inefficient, especially for spherical cameras with unfixed rotation angles and fields of view, which require manual calibration of multiple preset positions, resulting in low efficiency.
By acquiring real-world images, controlling the movement of the virtual-world image acquisition device based on the initial position and posture, automatically matching the virtual images and determining the projection matrix, automated camera calibration is achieved.
It achieves fast and automated camera calibration, improves the efficiency of camera calibration in the digital twin world, and reduces manual intervention and calibration time.
Smart Images

Figure CN120672860A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data calibration in large model data processing, and specifically to a calibration method, electronic device, storage medium and product of an image acquisition device. Background Art
[0002] Establishing a mapping relationship between the real world and the virtual world is the foundation of digital twins. Using various types of sensors, real-world target information can be mapped to the digital twin. Cameras, as a common sensor, are widely used. Mapping 2D image information captured by cameras to the 3D digital twin requires camera calibration technology.
[0003] Due to the special nature of digital twin camera calibration, traditional checkerboard methods and other automated calibration methods cannot be used. Current digital twin camera calibration technology can only be performed by manually clicking on 2D-3D sample pairs. While this can yield accurate results, it consumes a significant amount of manpower and time. For gun-type cameras (referred to as "gun cameras") that cannot rotate, a single manual calibration is sufficient. However, for spherical cameras (referred to as "dome cameras") with unfixed rotation angles and fields of view, several or even dozens of preset positions are often required, and calibration of all the preset position images is required to meet usage requirements. In the digital twin industry, which requires the use of hundreds or even thousands of cameras, this results in low efficiency when calibrating cameras.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0005] The embodiments of the present application provide a calibration method for an image acquisition device, an electronic device, a storage medium, and a product to at least solve the technical problem of low efficiency in camera calibration based on the digital twin world in the related art.
[0006] According to one aspect of an embodiment of the present application, a calibration method for an image acquisition device is provided, including: acquiring a real image acquired by a first image acquisition device in the real world; controlling the movement of a second image acquisition device in a virtual world based on an initial position and an initial posture of the first image acquisition device, and acquiring a set of virtual images acquired by the second image acquisition device during the movement, wherein the virtual world is a world obtained by simulating the real world; determining a target virtual image that successfully matches the real image from the set of virtual images; and determining a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and device parameters corresponding to the target virtual image, wherein the first projection matrix is used to characterize the mapping relationship between the virtual world and the real world.
[0007] According to another aspect of an embodiment of the present application, a calibration method for an image acquisition device is also provided, including: obtaining a real image by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes the real image, and the real image is acquired by the first image acquisition device in the real world; based on the initial position and initial posture of the first image acquisition device, controlling the movement of a second image acquisition device in the virtual world, and acquiring a set of virtual images acquired by the second image acquisition device during the movement, wherein the virtual world is a world obtained by simulating the real world; determining a target virtual image that successfully matches the real image from the set of virtual images; determining a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and the device parameters corresponding to the target virtual image, wherein the first projection matrix is used to characterize the mapping relationship between the virtual world and the real world; outputting the first projection matrix by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the first projection matrix.
[0008] According to another aspect of an embodiment of the present application, a calibration device for an image acquisition device is also provided, including: an acquisition module for acquiring a real image acquired by a first image acquisition device in the real world; a control module for controlling the movement of a second image acquisition device in a virtual world based on an initial position and an initial posture of the first image acquisition device, and acquiring a set of virtual images acquired by the second image acquisition device during the movement, wherein the virtual world is a world obtained by simulating the real world; a first determination module for determining a target virtual image that successfully matches the real image from the set of virtual images; and a second determination module for determining a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and device parameters corresponding to the target virtual image, wherein the first projection matrix is used to characterize the mapping relationship between the virtual world and the real world.
[0009] According to another aspect of an embodiment of the present application, a calibration device for an image acquisition device is provided, including: an acquisition module for acquiring a real image by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes a real image, and the real image is acquired by the first image acquisition device in the real world; a control module for controlling the movement of a second image acquisition device in a virtual world based on an initial position and an initial posture of the first image acquisition device, and acquiring a set of virtual images acquired by the second image acquisition device during the movement, wherein the virtual world is a world obtained by simulating the real world; a first determination module for determining a target virtual image that successfully matches the real image from the set of virtual images; a second determination module for determining a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and device parameters corresponding to the target virtual image, wherein the first projection matrix is used to represent a mapping relationship between the virtual world and the real world; and an output module for outputting the first projection matrix by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the first projection matrix.
[0010] According to another aspect of the embodiments of the present application, an electronic device is provided, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in the various embodiments of the present application when running.
[0011] According to another aspect of an embodiment of the present application, a computer-readable storage medium is also provided, which includes a stored executable program, wherein when the executable program is running, the device where the computer-readable storage medium is located is controlled to execute the methods in various embodiments of the present application.
[0012] According to another aspect of the embodiments of the present application, a computer program product is further provided, including a computer program, which implements the methods in various embodiments of the present application when executed by a processor.
[0013] According to another aspect of an embodiment of the present application, a computer program product is further provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method in each embodiment of the present application is implemented.
[0014] According to another aspect of the embodiments of the present application, a computer program is further provided, which implements the methods in various embodiments of the present application when executed by a processor.
[0015] In an embodiment of the present application, a method is employed to obtain a real image captured by a first image acquisition device in the real world; based on the initial position and initial posture of the first image acquisition device, control the movement of a second image acquisition device in the virtual world, and obtain a set of virtual images captured by the second image acquisition device during the movement, wherein the virtual world is a world simulated by the real world; determine a target virtual image that successfully matches the real image from the set of virtual images; and determine a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and device parameters corresponding to the target virtual image, wherein the first projection matrix is used to represent the mapping relationship between the virtual world and the real world. It is easy to notice that, based on the initial position and initial posture of the first image acquisition device, a set of virtual images of the second image acquisition device during the movement can be automatically obtained, and the target virtual image can be automatically determined from the set of virtual images, and finally the first projection matrix can be automatically obtained, without the need for manual calibration of a large amount of data, achieving the purpose of fast and automated camera calibration, thereby achieving the technical effect of efficient camera calibration based on the digital twin world, thereby solving the technical problem of low efficiency in camera calibration based on the digital twin world in the related art.
[0016] It is easy to notice that the above general description and the following detailed description are merely for the purpose of exemplifying and explaining the present application, and do not constitute a limitation of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a calibration method for an image acquisition device according to an embodiment of the present application;
[0019] Figure 2 is a structural block diagram of a computing environment according to an embodiment of the present application;
[0020] Figure 3 This is a structural block diagram of a service grid according to an embodiment of the present application;
[0021] Figure 4 is a flow chart of a calibration method for an image acquisition device according to Example 1 of the present application;
[0022] Figure 5 This is a schematic diagram of generating an optional homography matrix according to Example 1 of the present application;
[0023] Figure 6 is a flowchart of an optional camera calibration method according to Example 1 of the present application;
[0024] Figure 7 is a flow chart of a calibration method for an image acquisition device according to Example 2 of the present application;
[0025] Figure 8 is a schematic diagram of a calibration device for an image acquisition device according to Example 3 of the present application;
[0026] Figure 9 is a schematic diagram of a calibration device for an image acquisition device according to Example 4 of the present application;
[0027] Figure 10 This is a structural block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0030] First, some nouns or terms that appear in the description of the embodiments of the present application are subject to the following interpretations:
[0031] Position (x, y, z): refers to the coordinates of the camera in the world, usually longitude, latitude, and altitude.
[0032] Posture (h, p, r): refers to the camera's yaw angle, pitch angle, and roll angle.
[0033] fov: field of view, camera field of view angle.
[0034] P: Camera projection matrix, obtained through camera calibration.
[0035] H: Homography matrix, indicating the transformation relationship between corresponding points in two images.
[0036] SIFT: Scale-Invariant Feature Transform, a classic feature for describing images.
[0037] Example 1
[0038] According to an embodiment of the present application, a calibration method for an image acquisition device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0039] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 This is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a calibration method for an image acquisition device according to an embodiment of the present application. Figure 1 As shown, the computer terminal 10 (or mobile device) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a Universal Serial Bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0040] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0041] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the methods in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the methods in the above embodiments. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0042] The transmission device 106 is used to receive or send data via a network. A specific example of the network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0043] The display may be, for example, a touch screen liquid crystal display (LCD), which enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0044] Figure 1 The hardware structure block diagram shown can be used not only as an exemplary block diagram of the above-mentioned computer terminal 10 (or mobile device), but also as an exemplary block diagram of the above-mentioned server. In an optional embodiment, Figure 2 The block diagram shows the use of the above Figure 1The computer terminal 10 (or mobile device) is shown as an embodiment of a computing node in the computing environment 201 . Figure 2 is a structural block diagram of a computing environment according to an embodiment of the present application, such as Figure 2 As shown, computing environment 201 includes multiple computing nodes (e.g., servers) (illustrated as 210-1, 210-2, ...) running on a distributed network. Each computing node includes local processing and memory resources, and end user 202 can remotely run applications or store data in computing environment 201. Applications can be provided as multiple services 220-1, 220-2, 220-3, and 220-4 in computing environment 201, representing services "A," "D," "E," and "H," respectively.
[0045] End user 202 can provide and access services through a web browser or other software application on a client. In some embodiments, the provisioning and / or request of end user 202 can be provided to the ingress gateway 230. The ingress gateway 230 may include a corresponding agent to handle the provisioning and / or request for services (one or more services provided in the computing environment 201).
[0046] Services are provided or deployed based on various virtualization technologies supported by the computing environment 201. In some embodiments, services can be provided based on virtual machine (VM)-based virtualization, container-based virtualization, and / or similar methods. Virtual machine-based virtualization can be to simulate a real computer by initializing a virtual machine, executing programs and applications without directly contacting any actual hardware resources. While virtual machines virtualize machines, according to container-based virtualization, containers can be started to virtualize the entire operating system (OS) so that multiple workloads can run on a single operating system instance.
[0047] In an embodiment based on container virtualization, several containers of a service can be assembled into a Pod (e.g., a Kubernetes Pod). Figure 2 As shown, service 220-2 can be equipped with one or more Pods 240-1, 240-2, ..., 240-N (collectively, Pods). A Pod can include a proxy 245 and one or more containers 242-1, 242-2, ..., 242-M (collectively, containers). One or more containers in a Pod handle requests related to one or more corresponding functions of the service. Proxy 245 typically controls network functions related to the service, such as routing and load balancing. Other services can also be equipped with similar Pods.
[0048] During operation, executing a user request from the end user 202 may require calling one or more services in the computing environment 201, and executing one or more functions of a service may require calling one or more functions of another service. Figure 2 As shown, service “A” 220 - 1 receives a user request from end user 202 from ingress gateway 230 , service “A” 220 - 1 may call service “D” 220 - 2 , and service “D” 220 - 2 may request service “E” 220 - 3 to perform one or more functions.
[0049] This computing environment can be a cloud computing environment, where resource allocation is managed by the cloud service provider, allowing for feature development without having to worry about implementing, adjusting, or scaling servers. This computing environment allows developers to execute code in response to events without building or maintaining complex infrastructure. Services can be partitioned to perform a set of functions that can scale independently and automatically, rather than scaling a single hardware device to handle the potential load.
[0050] In another optional embodiment, Figure 3 The block diagram shows the use of the above Figure 1 The computer terminal 10 (or mobile device) is shown as an embodiment of the service grid. Figure 3 This is a structural diagram of a service grid according to an embodiment of the present application. Figure 3 As shown, the service grid 300 is mainly used to facilitate secure and reliable communication between multiple microservices. Microservices refer to decomposing an application into multiple smaller services or instances and distributing them to run on different clusters / machines.
[0051] like Figure 3 As shown, the microservices may include application service instance A and application service instance B, which form the functional application layer of the service grid 300. In one embodiment, application service instance A runs in the form of container / process 308 on machine / workload container group 314 (Pod), and application service instance B runs in the form of container / process 310 on machine / workload container group 316 (Pod).
[0052] In one implementation, application service instance A may be a product query service, and application service instance B may be a product ordering service.
[0053] like Figure 3As shown, application service instance A and grid proxy (sidecar) 303 coexist in machine workload container group 314, while application service instance B and grid proxy 305 coexist in machine workload container 316. Grid proxy 303 and grid proxy 305 form the data plane layer (dataplane) of service grid 300. Grid proxy 303 and grid proxy 305 run as container / process 304 and container / process 306, respectively, and can receive requests 312 for product query services. Bidirectional communication is possible between grid proxy 303 and application service instance A, and between grid proxy 305 and application service instance B. Furthermore, bidirectional communication is possible between grid proxy 303 and grid proxy 305.
[0054] In one embodiment, the traffic of application service instance A is routed to the appropriate destination via grid proxy 303, and the network traffic of application service instance B is routed to the appropriate destination via grid proxy 305. It should be noted that the network traffic mentioned herein includes, but is not limited to, Hypertext Transfer Protocol (HTTP), Representational State Transfer (REST), the high-performance, general-purpose open source framework (Google Remote Procedure Call, gRPC), the open source in-memory data structure storage system (Redis), and other forms.
[0055] In one embodiment, the data plane layer's functionality can be extended by writing custom filters for the proxy (Envoy) in service mesh 300. Service mesh proxy configuration can be designed to enable the service mesh to correctly proxy service traffic, enabling service interoperability and service governance. Mesh proxy 303 and mesh proxy 305 can be configured to perform at least one of the following functions: service discovery, health checking, routing, load balancing, authentication and authorization, and observability.
[0056] like Figure 3 As shown, the service grid 300 also includes a control plane layer. The control plane layer can be a group of services running in a dedicated namespace, and these services are hosted by a hosting control plane component 301 in a machine / workload container group (machine / Pod) 302. Figure 3 As shown, managed control plane component 301 communicates bidirectionally with mesh proxy 303 and mesh proxy 305. Managed control plane component 301 is configured to perform certain control and management functions. For example, managed control plane component 301 receives telemetry data transmitted by mesh proxy 303 and mesh proxy 305 and can further aggregate this telemetry data. In addition to these services, managed control plane component 301 can also provide a user-oriented application programming interface (API) to facilitate manipulation of network behavior and provide configuration data to mesh proxy 303 and mesh proxy 305.
[0057] Under the above operating environment, this application provides Figure 4 The calibration method of the image acquisition device shown. Figure 4 This is a flow chart of a calibration method for an image acquisition device according to Example 1 of the present application. Figure 4 As shown, it includes a client 40 and a cloud 42, wherein the client 40 and the cloud 42 are connected via a network. The client executes: sending a real image and receiving a first projection matrix. The cloud executes: obtaining a real image captured by a first image acquisition device in the real world; based on the initial position and initial posture of the first image acquisition device, controlling the movement of a second image acquisition device in the virtual world, and obtaining a set of virtual images captured by the second image acquisition device during the movement; determining a target virtual image that successfully matches the real image from the set of virtual images; and determining a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and the device parameters corresponding to the target virtual image. The method includes the following steps:
[0058] Step S402: Acquire a real image captured by a first image capture device in the real world.
[0059] The first image acquisition device may be a camera in the real world. The real image may be an image captured by the camera in the real world and containing real objects in the real world. For example, the real image may include, but is not limited to, real objects such as buildings, natural scenery, vehicles, and humans.
[0060] In an optional embodiment, when a mapping relationship between the real world and the virtual world needs to be established, a real image can first be captured by a first image capture device in the real world, and then the cloud can obtain the real image captured by the first image capture device through the network. For example, a real object in the real world can first be captured by a camera in the real world (i.e., the first image capture device) to obtain a real image, and then the cloud can obtain the real image through the network.
[0061] Step S404: Based on the initial position and initial posture of the first image acquisition device, control the movement of the second image acquisition device in the virtual world, and obtain a set of virtual images acquired by the second image acquisition device during the movement, wherein the virtual world is a world obtained by simulating the real world.
[0062] The above-mentioned initial position and initial posture are the initial installation position and initial installation posture of the first image acquisition device, wherein the initial position and initial posture can be determined based on the installation position information of the first image acquisition device and the anchor image in the digital twin world, or can be determined by an image database established offline. In the case of determining the initial position and initial posture based on the installation position information and anchor image of the first image acquisition device, the first image acquisition device generally provides rough installation position information when it is installed, so the image position of the real image can be determined; the anchor image is obtained by rotating the virtual camera in the digital twin world. Since the position and posture information corresponding to the anchor image is known, the initial position and initial information of the real image can be obtained by determining the anchor image whose matching degree with the real image is greater than the matching degree threshold. Furthermore, since the real image is acquired by the first image acquisition device, the initial position and initial posture of the real image can be regarded as the initial position and initial posture of the first image acquisition device. When determining the initial position and initial posture based on an offline image database, first, an anchor image whose distance from the camera screen of the first image acquisition device is less than a distance threshold can be retrieved, and secondly, the position and posture of the anchor image can be used as the initial position and initial posture of the first image acquisition device. The initial position may include but is not limited to: (x, y, z), where (x, y, z) is the coordinates of the first image acquisition device in the real world, x is longitude, y is latitude, and z is altitude. The initial posture may include but is not limited to: (h, p, r), where h is the yaw angle of the first image acquisition device in the real world, p is the pitch angle of the first image acquisition device in the real world, and r is the roll angle of the first image acquisition device in the real world. The above-mentioned virtual world is a digital twin world. The above-mentioned second image acquisition device can be a virtual camera in the virtual world.
[0063] In an optional embodiment, in a scene that does not include a field of view angle, when the initial position and initial posture of the first image acquisition device are obtained, the movement of the virtual camera (i.e., the second image acquisition device) in the virtual world can be controlled based on the initial position and initial posture to obtain a virtual image set. Among them, I i represents the i-th virtual image, and N represents the number of virtual preset positions.
[0064] For example, when the initial position and initial posture of the first image acquisition device are obtained, the virtual camera can be controlled to perform translation or rotation movement according to a preset interval based on the initial position and initial posture, so that the virtual camera can be moved to different virtual preset positions. Secondly, when the virtual camera moves to different virtual preset positions, a frame of virtual image I of the virtual camera at the virtual preset position can be obtained. i Then, the virtual camera can be controlled to move to the next virtual preset position according to the preset interval, and a frame of virtual image I of the virtual camera at the next virtual preset position can be obtained. i+1 , until the virtual camera moves through all virtual preset positions, the virtual images on N different virtual preset positions can be summarized, that is, the virtual image set can be obtained
[0065] In another optional embodiment, in a scene including a field of view, when the initial position and initial posture of the first image acquisition device are obtained, the movement of the virtual camera (i.e., the second image acquisition device) in the virtual world can be controlled based on the initial position, initial posture, and initial field of view FOV to obtain a virtual image set. Among them, I i represents the i-th virtual image, and N represents the number of virtual preset positions. It should be noted that because the virtual camera is located in the virtual world, when the virtual camera moves, it also needs to move according to the initial field of view in the virtual world. The angle of the initial field of view can be set to 60 degrees, but is not limited thereto. The specific value can be determined by the rendering engine and is not specifically limited in this embodiment. The virtual preset positions can be different virtual positions of the virtual camera in the virtual world set in advance by the user.
[0066] For example, when the initial position and initial posture of the first image acquisition device are obtained, the virtual camera can be controlled to translate, rotate, or change the field of view angle according to a preset interval based on the initial position, initial posture, and initial field of view. In this case, the virtual camera can be moved to different virtual preset positions. Secondly, after the virtual camera moves to different virtual preset positions, a frame of virtual image I of the virtual camera at the virtual preset position can be obtained. i Then, the virtual camera can be controlled to move to the next virtual preset position according to the preset interval, and a frame of virtual image I of the virtual camera at the next virtual preset position can be obtained. i+1 , until the virtual camera moves through all virtual preset positions, the virtual images on N different virtual preset positions can be summarized, that is, the virtual image set can be obtained
[0067] Optionally, when the virtual camera is translated according to the translation interval (dx, dy, dz), the position of the virtual preset position can be obtained as (x+dx, y+dy, z+dz). At this time, a frame of virtual image I of the virtual camera at the virtual preset position can be obtained through the twin digital engine. i Similarly, the virtual camera can be controlled to translate through all virtual preset positions in sequence according to the translation interval, and then a frame of virtual image of the virtual camera at all virtual preset positions can be obtained through the twin digital engine. Finally, all virtual images can be summarized, that is, a virtual image set can be obtained.
[0068] Optionally, when the virtual camera rotates according to the rotation interval (dh, dp, dr), the posture of the virtual preset position can be obtained as (h+dh, p+dp, r+dr). At this time, a frame of virtual image I of the virtual camera at the posture of the virtual preset position can be obtained through the twin digital engine. i Similarly, the virtual camera can be controlled to rotate through all the postures of the virtual preset position according to the rotation interval, and then a frame of virtual image of the virtual camera at different postures of the virtual preset position can be obtained through the twin digital engine. Finally, all virtual images can be summarized, that is, a virtual image set can be obtained.
[0069] Optionally, when the virtual camera is translated and rotated according to the translation interval (dx, dy, dz) and the rotation interval (dh, dp, dr), the position of the virtual preset position can be obtained as (x+dx, y+dy, z+dz), and the posture of the virtual preset position can be obtained as (h+dh, p+dp, r+dr). At this time, a frame of virtual image I can be obtained at the posture of the virtual position through the twin digital engine. i Similarly, the virtual camera can be controlled to translate and rotate through all virtual preset positions according to the translation interval and rotation interval, and then a frame of virtual image of the virtual camera in different postures at different virtual preset positions can be obtained through the twin digital engine. Finally, all virtual images can be summarized, that is, a virtual image set can be obtained.
[0070] Optionally, when the virtual camera is translated and the field of view angle is changed according to the translation interval (dx, dy, dz) and the field of view angle change interval dfov, the position of the virtual preset position can be obtained as (x+dx, y+dy, z+dz) and the field of view angle is (fov+dfov). At this time, a frame of virtual image I can be obtained at the field of view angle of the virtual position through the twin digital engine. iSimilarly, the virtual camera can be controlled to change the interval according to the translation interval and field of view angle. After the translation and field of view angle change through all virtual preset positions, a frame of virtual image of the virtual camera at different virtual preset positions and different field of view angles is obtained through the twin digital engine, and all virtual images are summarized to obtain a virtual image set.
[0071] Optionally, when the virtual camera rotates and changes the field of view according to the rotation interval (dh, dp, dr) and the field of view change interval dfov, the posture of the virtual preset position can be obtained as (h+dh, p+dp, r+dr), and the field of view is (fov+dfov). At this time, a frame of virtual image I can be obtained at the posture and field of view of the virtual position through the twin digital engine. i Similarly, the virtual camera can be controlled to change the interval according to the rotation interval and field of view angle. After the rotation and field of view angle change through all the postures and field of view angles of the virtual preset position, a frame of virtual image of the virtual camera at different postures and different field of view angles of the virtual preset position is obtained through the twin digital engine, and all virtual images are summarized to obtain a virtual image set.
[0072] Optionally, when the virtual camera is translated and rotated and the field of view angle is changed according to the translation interval (dx, dy, dz), the rotation interval (dh, dp, dr), and the field of view angle change interval dfov, the position of the virtual preset position can be obtained as (x+dx, y+dy, z+dz), the posture is (h+dh, p+dp, r+dr), and the field of view angle is (fov+dfov). At this time, a frame of virtual image I can be obtained at the posture and field of view angle of the virtual position through the twin digital engine. i Similarly, the virtual camera can be controlled to change the interval according to the translation interval, rotation interval and field of view angle. After the translation, rotation and field of view angle change through all virtual preset positions, a frame of virtual image of the virtual camera in different postures and different field of view angles at different virtual preset positions is obtained through the twin digital engine, and all virtual images are summarized to obtain a virtual image set.
[0073] Step S406: determining a target virtual image that successfully matches the real image from the virtual image set.
[0074] The target virtual image may be a virtual image in the virtual image set that successfully matches the real image and has the highest similarity.
[0075] It's important to note that virtual images follow different patterns than real images. Real images are created through optical imaging, while virtual images are created through computer rendering. The distribution of features like texture and color in virtual images differs significantly from those in real images. Therefore, cross-modal matching is necessary when matching virtual and real images.
[0076] In an optional embodiment, after obtaining the virtual image set, feature extraction can be performed on each virtual image in the virtual image set to obtain multiple virtual image features, and image feature extraction can be performed on the real image to obtain real image features. For example, a SIFT feature extractor can be used to perform SIFT feature extraction on each virtual image in the virtual image set to obtain multiple virtual SIFT features, and a SIFT feature extractor can be used to perform SIFT feature extraction on the real image to obtain real SIFT features. Next, the multiple virtual image features can be matched with the real image features to obtain multiple matching degrees, and the matching degrees can be sorted from largest to smallest. Finally, the virtual image corresponding to the virtual image feature with the highest matching degree can be used as the target virtual image. Alternatively, a feature matcher can be used to perform cross-modal matching on each virtual image feature and the real image feature, and the virtual image with the greatest number of matching points with the real image I can be used as the target virtual image. For example, a feature matcher can be used to perform matching calculations (i.e., cross-modal matching) on each virtual SIFT feature and the real SIFT feature, and the virtual image with the greatest number of matching points with the real image can be used as the target virtual image.
[0077] It's important to note that the SIFT feature extractor is a classic feature extractor. SIFT is a feature detection algorithm used in image processing and computer vision. SIFT features are invariant to image scaling, rotation, and certain perspective changes and brightness variations. Therefore, they are very effective in many vision applications, such as image registration, object recognition, 3D reconstruction, and robot navigation.
[0078] Step S408 : determining a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and device parameters corresponding to the target virtual image, wherein the first projection matrix is used to represent a mapping relationship between the virtual world and the real world.
[0079] The aforementioned device parameters may include, but are not limited to: a target position, a target posture, and a target field of view angle of the target virtual image in the virtual world.
[0080] In an optional embodiment, first, real image feature points of the real image can be obtained, and then the position information of the real image feature points can be used to deduce the internal and external parameters of the first image acquisition device, which may include but are not limited to: focal length, principal point position, distortion coefficient, etc. Then, according to the device parameters corresponding to the target virtual image, the internal and external parameters of the first image acquisition device are matched and calibrated with the parameters of the target virtual image, so that the first projection matrix of the first image acquisition device can be determined.
[0081] In another optional embodiment, first, the homography matrix H between the real image and the target virtual image can be determined based on the real image and the target virtual image, and then the second projection matrix P of the second image acquisition device can be determined based on the device parameters. b Finally, the product of the homography matrix and the second projection matrix can be obtained to obtain the first projection matrix P.
[0082] In an embodiment of the present application, a method is employed to obtain a real image captured by a first image acquisition device in the real world; based on the initial position and initial posture of the first image acquisition device, control the movement of a second image acquisition device in the virtual world, and obtain a set of virtual images captured by the second image acquisition device during the movement, wherein the virtual world is a world simulated by the real world; determine a target virtual image that successfully matches the real image from the set of virtual images; and determine a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and device parameters corresponding to the target virtual image, wherein the first projection matrix is used to represent the mapping relationship between the virtual world and the real world. It is easy to notice that, based on the initial position and initial posture of the first image acquisition device, a set of virtual images of the second image acquisition device during the movement can be automatically obtained, and the target virtual image can be automatically determined from the set of virtual images, and finally the first projection matrix can be automatically obtained, without the need for manual calibration of a large amount of data, achieving the purpose of fast and automated camera calibration, thereby achieving the technical effect of efficient camera calibration based on the digital twin world, thereby solving the technical problem of low efficiency in camera calibration based on the digital twin world in the related art.
[0083] In the above-mentioned embodiment of the present application, based on the initial position and initial posture of the first image acquisition device, the movement of the second image acquisition device in the virtual world is controlled, and a set of virtual images acquired by the second image acquisition device during the movement is obtained, including: controlling the second image acquisition device to be in the initial position, initial posture and preset field of view angle; controlling the second image acquisition device to move according to a preset movement interval so that the second image acquisition device moves to different target positions, wherein the movement includes at least one of the following: translation, rotation, and change of field of view angle; when the second image acquisition device moves to different target positions, acquiring the virtual image acquired by the second image acquisition device at the target position to obtain a set of virtual images.
[0084] The above-mentioned preset field of view angle is the initial field of view angle. The initial value of the preset field of view angle can be 60 degrees, but is not limited thereto. Among them, setting the preset field of view angle to 60 degrees gives the virtual world a wider field of view, which can enhance the realism and immersion of the virtual world. The above-mentioned preset movement interval can be a translation interval (dx, dy, dz), or a rotation interval (dh, dp, dr), or a field of view change interval dfov.
[0085] In an optional embodiment, after obtaining the initial position and initial posture of the first image acquisition device, the second image acquisition device can be controlled to be in the initial position, initial posture and preset field of view in the virtual world. Secondly, the second image acquisition device can be controlled to move according to the preset movement interval so that the second image acquisition device moves to different target positions (i.e., virtual preset positions). When the second image acquisition device moves to different target positions, the virtual images captured by the second image acquisition device at different target positions can be obtained. Finally, the different virtual images can be aggregated, that is, a virtual image set can be obtained.
[0086] Optionally, when the virtual camera is translated and the field of view angle is changed according to the translation interval (dx, dy, dz) and the field of view angle change interval dfov, the position of the virtual preset position can be obtained as (x+dx, y+dy, z+dz) and the field of view angle is (fov+dfov). At this time, a frame of virtual image I can be obtained at the field of view angle of the virtual position through the twin digital engine. i Similarly, the virtual camera can be controlled to change the interval according to the translation interval and field of view angle. After the translation and field of view angle change through all virtual preset positions, a frame of virtual image of the virtual camera at different virtual preset positions and different field of view angles is obtained through the twin digital engine, and all virtual images are summarized to obtain a virtual image set.
[0087] Optionally, when the virtual camera rotates and changes the field of view according to the rotation interval (dh, dp, dr) and the field of view change interval dfov, the posture of the virtual preset position can be obtained as (h+dh, p+dp, r+dr), and the field of view is (fov+dfov). At this time, a frame of virtual image I can be obtained at the posture and field of view of the virtual position through the twin digital engine. i Similarly, the virtual camera can be controlled to change the interval according to the rotation interval and field of view angle. After the rotation and field of view angle change through all the postures and field of view angles of the virtual preset position, a frame of virtual image of the virtual camera at different postures and different field of view angles of the virtual preset position is obtained through the twin digital engine, and all virtual images are summarized to obtain a virtual image set.
[0088] Optionally, when controlling the virtual camera to translate and rotate and change the field of view according to the translation interval (dx, dy, dz), the rotation interval (dh, dp, dr), and the field of view change interval dfov, firstly, the position of the virtual preset position can be obtained as (x+dx, y+dy, z+dz), the posture is (h+dh, p+dp, r+dr), and the field of view is (fov+dfov), and then a frame of virtual image I can be captured at the posture and field of view of the virtual position through the twin digital engine. i Similarly, the virtual camera can be controlled according to the translation interval, rotation interval and field of view interval. After the translation, rotation and field of view angle change through all virtual preset positions, a frame of virtual image of the virtual camera in different postures and different field of view angles at different virtual preset positions can be obtained through the twin digital engine, and the collected virtual images can be summarized to obtain a virtual image set.
[0089] In the above embodiment of the present application, the method further includes: obtaining installation position information of the first image acquisition device in the real world; and determining an initial position and an initial posture based on the real image and the installation position information.
[0090] The aforementioned installation position information may be the installation position information of the first image acquisition device, and may include but is not limited to: the installation position and installation posture of the first image acquisition device, wherein the installation position and installation posture are known values.
[0091] In an optional embodiment, the installation position information of the first image acquisition device in the real world can be obtained first. Secondly, the virtual camera in the virtual world can be rotated to obtain multiple anchor images. Then, the multiple anchor images can be matched with the real image. Since the position and posture information corresponding to the anchor images are known, the initial position and initial information of the real image can be obtained by determining the anchor images whose matching degree with the real image is greater than the matching degree threshold. Furthermore, since the real image is acquired by the first image acquisition device, the initial position and initial posture of the real image can be regarded as the initial position and initial posture of the first image acquisition device.
[0092] In the above embodiment of the present application, determining a target virtual image that successfully matches a real image from a virtual image set includes: using a scale-invariant feature transform extractor to extract features of the real image to obtain real image features, and using the scale-invariant feature transform extractor to extract features of each virtual image in the virtual image set to obtain virtual image features of each virtual image; matching the real image features with the virtual image features to determine the number of matching points between each virtual image in the virtual image set and the real image; and determining the virtual image corresponding to the maximum number of matching points as the target virtual image.
[0093] The scale-invariant feature transform extractor is a SIFT feature extractor. The real image features can be real SIFT features. The virtual image features can be virtual SIFT features.
[0094] In an optional embodiment, first, a SIFT feature extractor can be used to extract features from a real image to obtain real SIFT features. At the same time, a SIFT feature extractor can be used to extract features from each virtual image in a virtual image set to obtain multiple virtual SIFT features corresponding to each virtual image. Secondly, a feature matcher can be used to perform matching calculations on the real SIFT features and the multiple virtual SIFT features to determine the number of matching points between each virtual image and the real image. Finally, the virtual SIFT feature corresponding to the maximum number of matching points and the corresponding virtual image can be used as the target virtual image.
[0095] In the above embodiment of the present application, a first projection matrix of a first image acquisition device is determined based on a real image, a target virtual image, and device parameters corresponding to the target virtual image, including: determining a homography matrix between the real image and the target virtual image based on the real image and the target virtual image; determining a second projection matrix of a second image acquisition device based on the device parameters; and obtaining the product of the homography matrix and the second projection matrix to obtain the first projection matrix.
[0096] The second projection matrix mentioned above can be the virtual camera projection matrix P b . The homography matrix is an important concept used in computer vision and computer graphics. It is a 3x3 matrix used to describe the projective transformation from plane to plane. In image processing, the homography matrix is often used to map points on one image to corresponding points on another image, thereby achieving operations such as image registration, correction, and fusion. The homography matrix can be estimated from known corresponding point pairs and can then be used to achieve geometric transformation and reconstruction of the image.
[0097] In an optional embodiment, first, the homography matrix between the real image and the target virtual image can be determined based on the real image and the target virtual image. Secondly, the second projection matrix of the second image acquisition device can be determined based on the device parameters. Finally, the product of the homography matrix and the second projection matrix can be obtained to obtain the first projection matrix. For example, first, feature extraction can be performed on the real image through a feature extraction network to obtain real image features, wherein the dimension of the real image features is w*h*c, wherein w is the width of the real image features, h is the height, and c is the feature dimension, and w and h represent the spatial dimension of the real image features. Secondly, feature extraction can be performed on the target virtual image through a feature extraction network to obtain target virtual image features, wherein the dimension of the target virtual image features is the same as the dimension of the real image features, which will not be repeated here. Secondly, the real image features can be converted into multiple real feature vectors, and the target virtual image features can be converted into multiple target virtual feature vectors. Then, the real feature vectors and the virtual feature vectors can be matched through a feature matching network to generate a matching relationship graph. Finally, the homography matrix can be determined based on the matching relationship graph.
[0098] In the above embodiment of the present application, based on the real image and the target virtual image, the homography matrix between the real image and the target virtual image is determined, including: using a feature extraction network to extract features of the real image to obtain a real feature tensor, and using the feature extraction network to extract features of the target virtual image to obtain a virtual feature tensor; converting the real feature tensor into multiple real feature vectors, and converting the virtual feature tensor into multiple virtual feature vectors; using a feature matching network to match the real feature vector and the virtual feature vector to generate a matching relationship graph, wherein the matching relationship graph is used to represent the correspondence between the real feature vector and the virtual feature vector when the real feature vector and the virtual feature vector are successfully matched; based on the matching relationship graph, the homography matrix is determined.
[0099] The aforementioned real feature tensor is the real image feature. The aforementioned virtual image tensor is the virtual image feature. The aforementioned feature vector can be obtained by extracting the feature tensor into w*h feature vectors of dimension c along the spatial direction, or by taking the maximum value of different dimensions of the feature tensor, or by using a dimensionality reduction method, such as principal component analysis or independent component analysis, but is not limited thereto.
[0100] In an optional embodiment, Figure 5 is a schematic diagram of generating an optional homography matrix according to Example 1 of the present application, such as Figure 5 As shown, first, the real image and the target virtual image can be input into the feature extraction network respectively, and the feature extraction network extracts features from the real image and the target virtual image respectively to obtain a real feature tensor and a virtual feature tensor. Secondly, the real feature tensor can be converted into multiple real feature vectors, and the virtual feature tensor can be converted into multiple virtual feature vectors. Then, the multiple real feature vectors and the multiple virtual feature vectors can be input into the feature matching network respectively, and the feature matching network matches the real feature vectors and the virtual feature vectors to generate a matching relationship graph. Finally, the homography matrix can be calculated based on the matching relationship graph. Among them, the matching relationship graph is used to represent the corresponding relationship between the real feature vector and the virtual feature vector when the real feature vector and the virtual feature vector are successfully matched.
[0101] Optionally, the matching relationship graph is a set of triples {(p_v, p_r, score)} i Where p_v = (x_v, y_v) is a feature point in the virtual image, p_r = (x_r, y_r) is a feature point in the real image, and score is the matching score between p_v and p_r. Score is calculated as follows: the feature vector corresponding to p_v is v, the feature vector corresponding to p_r is r, and score is the cosine distance between v and r. The feature matching network outputs all possible matching triplets. By further filtering the scores using a threshold, we can obtain matching triplets with higher confidence (i.e., the matching relationship graph).
[0102] In another optional embodiment, the relationship between the homography matrix H and the matching point pairs is:
[0103]
[0104] By matching the relationship graph, several pairs (no less than 4 pairs) of points in the virtual image and the real image that are in a corresponding relationship can be obtained. Among them, there are many classic methods that can be used to solve the homography matrix, such as the direct linear method, which is not specifically limited in this embodiment.
[0105] It should be noted that the feature extraction network in this embodiment is obtained by fine-tuning and training on a general feature extraction network, which can focus more on the key point information of the image. The second projection matrix mentioned above can be the virtual camera projection matrix P b .
[0106] In the above embodiment of the present application, a feature matching network is used to match the real feature vector and the virtual feature vector to obtain a matching relationship diagram, including: using the feature matching network to match the real feature vector and the virtual feature vector to obtain a matching score between the real feature vector and the virtual feature vector; when the matching score is greater than a preset score, it is determined that the real feature vector and the virtual feature vector are successfully matched, and a matching relationship diagram is generated based on the real feature vector, the virtual feature vector and the matching score.
[0107] The preset score can be set in advance by the user and is used to determine whether the real feature vector and the virtual feature vector are matched successfully. If the matching score is greater than the preset score, it can be determined that the real feature vector and the virtual feature vector are matched successfully. If the matching score is less than or equal to the preset score, it can be determined that the real feature vector and the virtual feature vector are not matched successfully.
[0108] In an optional embodiment, the real feature vector and the virtual feature vector can first be matched through a feature matching network to obtain a matching score between the real feature vector and the virtual feature vector. Secondly, the matching score can be compared with a preset score. When the matching score is greater than the preset score, it can be determined that the real feature vector and the virtual feature vector are matched successfully. At this time, a matching relationship graph can be generated based on the real feature vector, the virtual feature vector and the matching score.
[0109] In the above embodiment of the present application, the real feature tensor is converted into multiple real feature vectors, and the virtual feature tensor is converted into multiple virtual feature vectors, including: expanding the real feature tensor according to the spatial dimension to obtain multiple real feature vectors; expanding the virtual feature tensor according to the spatial dimension to obtain multiple virtual feature vectors.
[0110] In an optional embodiment, the real feature tensor can be first expanded based on the spatial dimension c, that is, w*h real feature vectors with a dimension of c can be obtained. At the same time, the virtual feature tensor can also be expanded based on the spatial dimension c, that is, w*h virtual feature vectors with a dimension of c can be obtained.
[0111] In the above embodiment of the present application, the method also includes: outputting an initial position and an initial posture; in response to an adjustment instruction for adjusting the initial position and the initial posture, determining a target position and a target posture corresponding to the adjustment instruction; and controlling the movement of the second image acquisition device based on the target position and the target posture.
[0112] In an optional embodiment, after obtaining the initial position and initial posture, the cloud can also send the initial position and initial posture to the client, and the initial position and initial posture are displayed to the user by the display interface of the client, so that the user can confirm the initial position and initial posture. If the user determines that the initial position and initial posture are inaccurate, the user can perform an adjustment operation on the input area in the display interface. At this time, the display interface can generate an adjustment instruction based on the received adjustment operation, and determine the corresponding target position and target posture based on the adjustment instruction, and then send the target position and target posture to the cloud, so that the cloud can control the movement of the second image acquisition device based on the target position and target posture.
[0113] In the above-mentioned embodiment of the present application, a first projection matrix of a first image acquisition device is determined based on a real image, a target virtual image, and device parameters corresponding to the target virtual image, including: outputting a virtual image set and a target virtual image; in response to a selection instruction for selecting other virtual images in the virtual image set, determining a virtual image corresponding to the selection instruction, wherein the other virtual images are virtual images other than the target virtual image in the virtual image set; and determining the first projection matrix based on the real image, the virtual image corresponding to the selection instruction, and device parameters corresponding to the virtual image corresponding to the selection instruction.
[0114] In an optional embodiment, after obtaining the virtual image set and the target virtual image, the cloud can send the virtual image set and the target virtual image to the client. The client can display the virtual image set and the target virtual image to the user through the display interface. When the user adjusts the target virtual image, the user can perform a selection operation on other virtual images in the virtual image set. At this time, the display interface of the client can generate a selection instruction based on the received selection operation, and determine that the virtual image corresponding to the selection instruction is the new target virtual image. At this time, the client can send the new target virtual image to the cloud, so that the cloud can subsequently determine the first projection matrix based on the real image, the virtual image corresponding to the selection instruction, and the device parameters corresponding to the virtual image corresponding to the selection instruction.
[0115] The present application embodiment develops and builds an automated calibration system based on the digital twin world. This system uses virtual image preset position selection technology and cross-modal image key point matching technology to match key points between the virtual image of the digital twin world with camera parameters and the real image captured by the camera in real time, eliminating the reliance on manual annotation and achieving automated generation of calibration results.
[0116] The embodiment of the present application is based on digital twins and builds a process framework that can automatically calibrate real images. First, the position and posture of the camera can be roughly determined in combination with the camera picture and camera position information, and the standardized preset position acquisition program can be called at this position to obtain multiple preset position pictures and filter out the target preset position pictures and their camera parameters. The virtual image is matched with the real image across modalities, and the homography matrix is calculated. Finally, the camera parameters of the real image can be determined based on the camera parameters of the virtual image preset position to complete the camera calibration. This framework can automatically complete the camera calibration in the digital twin scenario, get rid of the dependence on manual annotation, and make it possible to quickly and scalably establish a mapping relationship from the real world to the virtual world.
[0117] Figure 6 is a flow chart of an optional camera calibration method according to Example 1 of the present application, such as Figure 6 As shown, the method may include the following steps:
[0118] Step S61: roughly determine the camera position and posture.
[0119] Specifically, the camera's location and orientation can be roughly determined based on the image content and pre-acquired user information. In the digital twin engine, the corresponding parameters can be set in advance to orient the virtual camera toward that location.
[0120] Step S62: acquiring a virtual preset position image.
[0121] Specifically, the camera is rotated and translated with the position (x, y, z), attitude (h, p, r), and field of view angle fov determined in step S61 as the center. Among them, the translation interval of the camera position is (dx, dy, dz), the rotation interval of the camera attitude is (dh, dp, dr), and the field of view angle interval is dfov. The camera position of the virtual preset position is (x+dx, y+dy, z+dz), the attitude parameters are (h+dh, p+dp, r+dr), and the field of view angle is fov+dfov. Using the digital twin engine, a frame of image I is obtained at each virtual preset position. i , the obtained virtual image set is recorded as Where N is the number of preset positions.
[0122] Step S63 , screening the target preset position (ie the position of the target virtual image) and acquiring parameters.
[0123] Specifically, the acquired virtual image set Match it with the image to be calibrated (i.e., the real image) I. Since the number of candidate preset positions of the virtual image is large, the SIFT feature extractor can be used to calibrate the virtual image I iThe SIFT features of the image to be calibrated are extracted respectively, and the feature matcher is used to perform matching calculations. The preset position image with the most matching points with the image to be calibrated I is the target preset position image I b .
[0124] Step S64: cross-modal matching of the virtual image and the real image.
[0125] Specifically, in order to accurately calculate the image to be calibrated I and the target preset position image (ie, the target virtual image) I b The homography matrix can be used to accurately calculate the homography matrix H of the two images through the cross-modal keypoint matching network. Specifically, i) first use the feature extraction network to extract the feature tensors of the virtual image and the real image; ii) each feature tensor is pulled into several feature vectors according to the position; iii) the feature matching network is used to match the feature vectors, and the matching score exceeding the threshold is considered a successful match, and the matching relationship graph is calculated; iv) the image to be calibrated I and the target preset position image I are calculated based on the matching relationship. b The homography matrix H of .
[0126] Step S65: generating calibration results.
[0127] Specifically, according to the camera parameters of the target matching preset position, its camera projection matrix P is calculated b . Match the homography matrix H with the camera projection matrix P of the target preset position b Multiply them together to get the camera projection matrix P of the current screen (P = H·P b ). The real image camera projection matrix P describes the mapping relationship between the 3D world and the 2D world, which is the calibration relationship matrix required by the embodiment of this application.
[0128] The cross-modal key point matching network proposed in the embodiment of the present application can stably match feature points of cross-modal virtual images and real images, and calculate the homography matrix H.
[0129] The embodiment of this application develops and builds an automated calibration system based on the digital twin world. This system gets rid of the need for manual labeling and realizes the fully automated generation of calibration results.
[0130] The cross-modal image key point matching technology of the embodiment of the present application matches the key points of the virtual image with camera parameters in the digital twin world with the real image acquired by the camera in real time.
[0131] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0132] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0133] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0134] Example 2
[0135] Figure 7 This is a flow chart of a calibration method for an image acquisition device according to Example 2 of the present application. Figure 7 As shown, the method includes the following steps:
[0136] Step S702: acquiring a real image by calling a first interface, wherein the first interface includes a first parameter, a parameter value of the first parameter includes a real image, and the real image is acquired by a first image acquisition device in the real world;
[0137] Step S704: Based on the initial position and initial posture of the first image acquisition device, controlling the movement of a second image acquisition device in the virtual world, and acquiring a set of virtual images captured by the second image acquisition device during the movement, wherein the virtual world is a world obtained by simulating the real world;
[0138] Step S706 , determining a target virtual image that successfully matches the real image from the virtual image set;
[0139] Step S708: determining a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and device parameters corresponding to the target virtual image, wherein the first projection matrix is used to represent a mapping relationship between the virtual world and the real world;
[0140] Step S7010: Output the first projection matrix by calling the second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the first projection matrix.
[0141] The first interface may be an interface for acquiring a real image from a client. The second interface may be an interface for sending a first projection matrix to a client.
[0142] In an optional embodiment, the cloud can first obtain a real image from the client through the first interface. Secondly, the cloud can control the movement of the second image acquisition device in the virtual world based on the initial position and initial posture of the first image acquisition device, and obtain a set of virtual images acquired by the second image acquisition device during the movement. Then, the cloud can determine the target virtual image that successfully matches the real image from the virtual image set. Then, the cloud can also determine the first projection matrix of the first image acquisition device based on the real image, the target virtual image, and the device parameters corresponding to the target virtual image. Finally, the cloud can output the first projection matrix to the client by calling the second interface.
[0143] Among them, the first interface includes a first parameter, the parameter value of the first parameter includes a real image, and the real image is captured by a first image acquisition device in the real world; the virtual world is a world obtained by simulating the real world; the first projection matrix is used to represent the mapping relationship between the virtual world and the real world; the second interface includes a second parameter, and the parameter value of the second parameter includes the first projection matrix.
[0144] Example 3
[0145] According to an embodiment of the present application, a calibration device for an image acquisition device is also provided for implementing the calibration method for the above-mentioned image acquisition device. Figure 8 is a schematic diagram of a calibration device for an image acquisition device according to Example 3 of the present application, such as Figure 8 As shown, the device includes: an acquisition module 82 , a control module 84 , a first determination module 86 and a second determination module 88 .
[0146] Among them, the acquisition module is used to obtain the real image captured by the first image acquisition device in the real world; the control module is used to control the movement of the second image acquisition device in the virtual world based on the initial position and initial posture of the first image acquisition device, and obtain the virtual image set captured by the second image acquisition device during the movement, wherein the virtual world is a world obtained by simulating the real world; the first determination module is used to determine the target virtual image that successfully matches the real image from the virtual image set; the second determination module is used to determine the first projection matrix of the first image acquisition device based on the real image, the target virtual image, and the device parameters corresponding to the target virtual image, wherein the first projection matrix is used to represent the mapping relationship between the virtual world and the real world.
[0147] It should be noted that the acquisition module 82, control module 84, first determination module 86, and second determination module 88 correspond to steps S402 to S408 in Example 1. The four modules and corresponding steps implement the same examples and application scenarios, but are not limited to the contents disclosed in Example 1. It should be noted that the modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The modules can also be part of a device and can be run in the computer terminal 10 provided in Example 1.
[0148] In the above embodiments of the present application, the control module includes: a first control unit, a second control unit, a first acquisition unit and a storage unit.
[0149] Among them, the first control unit is used to control the second image acquisition device to be in an initial position, initial posture and preset field of view angle; the second control unit is used to control the second image acquisition device to move according to a preset movement interval so that the second image acquisition device moves to different target positions, wherein the movement includes at least one of the following: translation, rotation, and change of field of view angle; the first acquisition unit is used to acquire the virtual image captured by the second image acquisition device at the target position when the second image acquisition device moves to different target positions, thereby obtaining a virtual image set.
[0150] In the above embodiment of the present application, the device further includes: a location information acquisition module and a third determination module.
[0151] Among them, the position information acquisition module is used to obtain the installation position information of the first image acquisition device in the real world; the third determination module is used to determine the initial position and initial posture based on the real image and the installation position information.
[0152] In the above embodiment of the present application, the first determination module includes: a feature extraction unit, a matching unit and a first determination unit.
[0153] Among them, the feature extraction unit is used to use a scale-invariant feature transform extractor to extract features from a real image to obtain real image features, and use the scale-invariant feature transform extractor to extract features from each virtual image in a virtual image set to obtain virtual image features of each virtual image; the matching unit is used to match real image features with virtual image features to determine the number of matching points between the virtual images in the virtual image set and the real images; the first determination unit is used to determine that the virtual image corresponding to the maximum number of matching points is the target virtual image.
[0154] In the above embodiment of the present application, the second determination module includes: a second determination unit, a third determination unit and a second acquisition unit.
[0155] Among them, the second determination unit is used to determine the homography matrix between the real image and the target virtual image based on the real image and the target virtual image; the third determination unit is used to determine the second projection matrix of the second image acquisition device based on the device parameters; and the second acquisition unit is used to obtain the product of the homography matrix and the second projection matrix to obtain the first projection matrix.
[0156] In the above embodiment of the present application, the second determination unit includes: a feature extraction subunit, a conversion subunit and a matching subunit.
[0157] Among them, the feature extraction subunit is used to use the feature extraction network to extract features of the real image to obtain a real feature tensor, and use the feature extraction network to extract features of the target virtual image to obtain a virtual feature tensor; the conversion subunit is used to convert the real feature tensor into multiple real feature vectors, and convert the virtual feature tensor into multiple virtual feature vectors; the matching subunit is used to use the feature matching network to match the real feature vector and the virtual feature vector to generate a matching relationship graph, wherein the matching relationship graph is used to represent the correspondence between the real feature vector and the virtual feature vector when the real feature vector and the virtual feature vector are successfully matched; based on the matching relationship graph, the homography matrix is determined.
[0158] In the above embodiment of the present application, the matching subunit is also used to: use the feature matching network to match the real feature vector and the virtual feature vector to obtain the matching score between the real feature vector and the virtual feature vector; when the matching score is greater than the preset score, it is determined that the real feature vector and the virtual feature vector are successfully matched, and a matching relationship diagram is generated based on the real feature vector, the virtual feature vector and the matching score.
[0159] In the above embodiment of the present application, the conversion subunit is also used to: expand the real feature tensor according to the spatial dimension to obtain multiple real feature vectors; expand the virtual feature tensor according to the spatial dimension to obtain multiple virtual feature vectors.
[0160] In the above embodiment of the present application, the apparatus further includes: a first output module, a fourth determination module and an image acquisition device control module.
[0161] Among them, the output module is used to output the initial position and initial posture; the fourth determination module is used to respond to the adjustment instruction for adjusting the initial position and initial posture, and determine the target position and target posture corresponding to the adjustment instruction; the image acquisition device control module is used to control the movement of the second image acquisition device based on the target position and target posture.
[0162] In the above embodiment of the present application, the second determination module further includes: an output unit, a fourth determination unit and a fifth determination unit.
[0163] The output unit is used to output a virtual image set and a target virtual image; the fourth determination unit is used to determine the virtual image corresponding to the selection instruction in response to a selection instruction for selecting other virtual images in the virtual image set, wherein the other virtual images are virtual images in the virtual image set other than the target virtual image; and the fifth determination unit is used to determine the first projection matrix based on the real image, the virtual image corresponding to the selection instruction, and the device parameters corresponding to the virtual image corresponding to the selection instruction.
[0164] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0165] Example 4
[0166] According to an embodiment of the present application, a calibration device for an image acquisition device is also provided for implementing the calibration method for the above-mentioned image acquisition device. Figure 9 is a schematic diagram of a calibration device for an image acquisition device according to Example 4 of the present application, such as Figure 9 As shown, the device includes: an acquisition module 92 , a control module 94 , a first determination module 96 , a second determination module 98 and an output module 910 .
[0167] Among them, the acquisition module is used to obtain a real image by calling a first interface, wherein the first interface includes a first parameter, and the parameter value of the first parameter includes a real image, and the real image is collected by a first image acquisition device in the real world; the control module is used to control the movement of a second image acquisition device in the virtual world based on the initial position and initial posture of the first image acquisition device, and obtain a set of virtual images collected by the second image acquisition device during the movement, wherein the virtual world is a world obtained by simulating the real world; the first determination module is used to determine a target virtual image that successfully matches the real image from the virtual image set; the second determination module is used to determine a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and the device parameters corresponding to the target virtual image, wherein the first projection matrix is used to represent the mapping relationship between the virtual world and the real world; the output module is used to output the first projection matrix by calling a second interface, wherein the second interface includes a second parameter, and the parameter value of the second parameter includes the first projection matrix.
[0168] It should be noted that the acquisition module 92, control module 94, first determination module 96, second determination module 98, and output module 910 described above correspond to steps S702 to S7010 in Example 2. The examples and application scenarios implemented by the five modules and the corresponding steps are the same, but are not limited to the contents disclosed in Example 1. It should be noted that the modules or units described above can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The modules can also be part of a device and can be run in the computer terminal 10 provided in Example 1.
[0169] It should be noted that the preferred implementation scheme involved in the above embodiments of this application is the same as the scheme provided in Example 1, as well as the application scenario and implementation process, but is not limited to the scheme provided in Example 1.
[0170] Example 5
[0171] The embodiment of the present application may provide an electronic device, which may be any electronic device in a group of electronic devices. Optionally, in this embodiment, the electronic device may also be replaced by a terminal device such as a mobile terminal.
[0172] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0173] In this embodiment, the computer terminal can execute the program code in the method.
[0174] Optionally, Figure 10This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 10 As shown, the electronic device A may include: one or more (only one is shown in the figure) processors 1002, a memory 1004, a storage controller, and a peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module and a display.
[0175] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implementing the methods in the above embodiments. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely located relative to the processor, and these remote memories may be connected to terminal A via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0176] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: obtain a real image captured by a first image acquisition device in the real world; based on the initial position and initial posture of the first image acquisition device, control the movement of a second image acquisition device in the virtual world, and obtain a set of virtual images captured by the second image acquisition device during the movement, wherein the virtual world is a world obtained by simulating the real world; determine a target virtual image that successfully matches the real image from the virtual image set; based on the real image, the target virtual image, and the device parameters corresponding to the target virtual image, determine the first projection matrix of the first image acquisition device, wherein the first projection matrix is used to represent the mapping relationship between the virtual world and the real world.
[0177] Optionally, the processor may also execute program codes for the following steps: controlling the second image acquisition device to be in an initial position, initial posture, and preset field of view angle; controlling the second image acquisition device to move according to a preset movement interval so that the second image acquisition device moves to different target positions, wherein the movement includes at least one of the following: translation, rotation, and change of field of view angle; when the second image acquisition device moves to different target positions, obtaining a virtual image captured by the second image acquisition device at the target position to obtain a virtual image set.
[0178] Optionally, the processor may further execute program codes of the following steps: obtaining installation position information of the first image acquisition device in the real world; and determining an initial position and an initial posture based on the real image and the installation position information.
[0179] Optionally, the processor may also execute the program code of the following steps: performing feature extraction on a real image using a scale-invariant feature transform extractor to obtain real image features, and performing feature extraction on each virtual image in a virtual image set using a scale-invariant feature transform extractor to obtain virtual image features of each virtual image; matching the real image features with the virtual image features to determine the number of matching points between each virtual image in the virtual image set and the real image; and determining the virtual image corresponding to the maximum number of matching points as the target virtual image.
[0180] Optionally, the processor may also execute the program code of the following steps: determining the homography matrix between the real image and the target virtual image based on the real image and the target virtual image; determining the second projection matrix of the second image acquisition device based on the device parameters; and obtaining the product of the homography matrix and the second projection matrix to obtain the first projection matrix.
[0181] Optionally, the processor may also execute the program code of the following steps: using a feature extraction network to extract features from a real image to obtain a real feature tensor, and using the feature extraction network to extract features from a target virtual image to obtain a virtual feature tensor; converting the real feature tensor into multiple real feature vectors, and converting the virtual feature tensor into multiple virtual feature vectors; using a feature matching network to match the real feature vector and the virtual feature vector to generate a matching relationship graph, wherein the matching relationship graph is used to characterize the correspondence between the real feature vector and the virtual feature vector when the real feature vector and the virtual feature vector are successfully matched; and determining the homography matrix based on the matching relationship graph.
[0182] Optionally, the processor may also execute the program code of the following steps: using a feature matching network to match the real feature vector and the virtual feature vector to obtain a matching score between the real feature vector and the virtual feature vector; when the matching score is greater than a preset score, determining that the real feature vector and the virtual feature vector are successfully matched, and generating a matching relationship diagram based on the real feature vector, the virtual feature vector and the matching score.
[0183] Optionally, the processor may further execute program code of the following steps: expanding the real feature vector according to the spatial dimension to obtain multiple real feature vectors; and expanding the virtual feature tensor according to the spatial dimension to obtain multiple virtual feature vectors.
[0184] Optionally, the processor may also execute program code for the following steps: outputting an initial position and an initial posture; in response to an adjustment instruction for adjusting the initial position and the initial posture, determining a target position and a target posture corresponding to the adjustment instruction; and controlling the movement of the second image acquisition device based on the target position and the target posture.
[0185] Optionally, the processor may also execute program code for the following steps: outputting a virtual image set and a target virtual image; determining, in response to a selection instruction for selecting other virtual images in the virtual image set, a virtual image corresponding to the selection instruction, wherein the other virtual images are virtual images in the virtual image set other than the target virtual image; and determining a first projection matrix based on the real image, the virtual image corresponding to the selection instruction, and device parameters corresponding to the virtual image corresponding to the selection instruction.
[0186] According to an embodiment of the present application, a method is provided for obtaining a real image captured by a first image acquisition device in the real world; controlling the movement of a second image acquisition device in the virtual world based on the initial position and initial posture of the first image acquisition device, and obtaining a set of virtual images captured by the second image acquisition device during the movement, wherein the virtual world is a world obtained by simulating the real world; determining a target virtual image that successfully matches the real image from the set of virtual images; and determining a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and device parameters corresponding to the target virtual image, wherein the first projection matrix is used to represent the mapping relationship between the virtual world and the real world. It is easy to notice that, based on the initial position and initial posture of the first image acquisition device, a set of virtual images of the second image acquisition device during the movement can be automatically obtained, and the target virtual image can be automatically determined from the set of virtual images, and finally the first projection matrix can be automatically obtained, without the need for manual calibration of a large amount of data, achieving the purpose of fast and automated camera calibration, thereby achieving the technical effect of efficient camera calibration based on the digital twin world, thereby solving the technical problem of low efficiency in camera calibration based on the digital twin world in the related art.
[0187] Those skilled in the art will understand that Figure 10 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, and other terminal devices. Figure 10 It does not limit the structure of the above electronic device. For example, the electronic device A may also include Figure 10 More or fewer components (such as network interfaces, display devices, etc.) shown in the figure, or having Figure 10 Different configurations shown.
[0188] A person skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0189] Example 6
[0190] The embodiment of the present application further provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the method provided in the above embodiment.
[0191] Optionally, in this embodiment, the storage medium may be located in any electronic device in a group of electronic devices in a computer network, or in any mobile terminal in a group of mobile terminals.
[0192] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: obtaining a real image captured by a first image capture device in the real world; controlling the movement of a second image capture device in the virtual world based on an initial position and an initial posture of the first image capture device, and obtaining a set of virtual images captured by the second image capture device during the movement, wherein the virtual world is a world obtained by simulating the real world; determining a target virtual image that successfully matches the real image from the set of virtual images; determining a first projection matrix of the first image capture device based on the real image, the target virtual image, and device parameters corresponding to the target virtual image, wherein the first projection matrix is used to characterize the mapping relationship between the virtual world and the real world.
[0193] Optionally, the computer-readable storage medium is further configured to store program code for executing the following steps: controlling the second image acquisition device to be in an initial position, an initial posture and a preset field of view angle; controlling the second image acquisition device to move according to a preset movement interval so that the second image acquisition device moves to different target positions, wherein the movement includes at least one of the following: translation, rotation, and change of field of view angle; when the second image acquisition device moves to different target positions, obtaining a virtual image captured by the second image acquisition device at the target position to obtain a virtual image set.
[0194] Optionally, the computer-readable storage medium is further configured to store program code for executing the following steps: obtaining installation position information of the first image acquisition device in the real world; and determining an initial position and an initial posture based on the real image and the installation position information.
[0195] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: using a scale-invariant feature transform extractor to extract features from a real image to obtain real image features, and using a scale-invariant feature transform extractor to extract features from each virtual image in a virtual image set to obtain virtual image features of each virtual image; matching the real image features with the virtual image features to determine the number of matching points between each virtual image in the virtual image set and the real image; and determining the virtual image corresponding to the maximum number of matching points as the target virtual image.
[0196] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: determining a homography matrix between the real image and the target virtual image based on the real image and the target virtual image; determining a second projection matrix of the second image acquisition device based on device parameters; and obtaining the product of the homography matrix and the second projection matrix to obtain a first projection matrix.
[0197] Optionally, the computer-readable storage medium is also configured to store program code for executing the following steps: using a feature extraction network to extract features from a real image to obtain a real feature tensor, and using the feature extraction network to extract features from a target virtual image to obtain a virtual feature tensor; converting the real feature tensor into multiple real feature vectors, and converting the virtual feature tensor into multiple virtual feature vectors; using a feature matching network to match the real feature vector and the virtual feature vector to generate a matching relationship graph, wherein the matching relationship graph is used to characterize the correspondence between the real feature vector and the virtual feature vector when the real feature vector and the virtual feature vector are successfully matched; and determining the homography matrix based on the matching relationship graph.
[0198] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: using a feature matching network to match the real feature vector and the virtual feature vector to obtain a matching score between the real feature vector and the virtual feature vector; when the matching score is greater than a preset score, determining that the real feature vector and the virtual feature vector are successfully matched, and generating a matching relationship graph based on the real feature vector, the virtual feature vector and the matching score.
[0199] Optionally, the computer-readable storage medium is further configured to store program code for executing the following steps: expanding the real feature vector according to the spatial dimension to obtain multiple real feature vectors; expanding the virtual feature tensor according to the spatial dimension to obtain multiple virtual feature vectors.
[0200] Optionally, the computer-readable storage medium is also configured to store program code for performing the following steps: outputting an initial position and an initial posture; in response to an adjustment instruction to adjust the initial position and the initial posture, determining a target position and a target posture corresponding to the adjustment instruction; and controlling the movement of the second image acquisition device based on the target position and the target posture.
[0201] Optionally, the computer-readable storage medium is further configured to store program code for executing the following steps: outputting a virtual image set and a target virtual image; in response to a selection instruction for selecting other virtual images in the virtual image set, determining a virtual image corresponding to the selection instruction, wherein the other virtual images are virtual images in the virtual image set other than the target virtual image; and determining a first projection matrix based on the real image, the virtual image corresponding to the selection instruction, and device parameters corresponding to the virtual image corresponding to the selection instruction.
[0202] Example 7
[0203] The embodiment of the present application further provides a computer program product. Optionally, in this embodiment, the computer program product may include a computer program, and when the computer program is executed by a processor, the method provided in the embodiment is implemented.
[0204] Example 8
[0205] The embodiments of the present application further provide a computer program product. Optionally, the computer program product may include a non-volatile computer-readable storage medium, which may be used to store a computer program that, when executed by a processor, implements the method provided in the embodiments above.
[0206] Example 9
[0207] The embodiment of the present application further provides a computer program. Optionally, in this embodiment, when the computer program is executed by a processor, the method provided in the above embodiment is implemented.
[0208] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0209] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0210] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0211] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0212] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0213] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0214] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A calibration method for an image acquisition device, characterized in that: include: Acquire a real image captured by a first image capture device in the real world; Based on the initial position and initial posture of the first image acquisition device, controlling the movement of a second image acquisition device in a virtual world, and obtaining a set of virtual images captured by the second image acquisition device during the movement, wherein the virtual world is a world obtained by simulating the real world; Determining a target virtual image that successfully matches the real image from the virtual image set; Based on the real image, the target virtual image, and device parameters corresponding to the target virtual image, a first projection matrix of the first image acquisition device is determined, wherein the first projection matrix is used to represent a mapping relationship between the virtual world and the real world.
2. The method according to claim 1, characterized in that The controlling the movement of the second image acquisition device in the virtual world based on the initial position and initial posture of the first image acquisition device, and obtaining a set of virtual images acquired by the second image acquisition device during the movement, includes: Controlling the second image acquisition device to be in the initial position, the initial posture and the preset field of view angle; Controlling the second image acquisition device to move according to a preset movement interval so that the second image acquisition device moves to different target positions, wherein the movement includes at least one of the following: translation, rotation, and change of field of view angle; When the second image acquisition device moves to the target position, the virtual image acquired by the second image acquisition device at the target position is acquired to obtain the virtual image set.
3. The method according to claim 1, characterized in that The method further comprises: Acquiring installation position information of the first image acquisition device in the real world; The initial position and the initial posture are determined based on the real image and the installation position information.
4. The method according to claim 1, wherein The determining, from the set of virtual images, a target virtual image that successfully matches the real image comprises: Performing feature extraction on the real image using a scale-invariant feature transform extractor to obtain real image features, and performing feature extraction on each virtual image in the virtual image set using the scale-invariant feature transform extractor to obtain virtual image features of each virtual image; Matching the real image features with each virtual image feature to determine the number of matching points between each virtual image in the virtual image set and the real image; The virtual image corresponding to the maximum number of matching points is determined as the target virtual image.
5. The method according to claim 1, wherein The determining of a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and device parameters corresponding to the target virtual image includes: determining a homography matrix between the real image and the target virtual image based on the real image and the target virtual image; determining a second projection matrix of the second image acquisition device based on the device parameters; Obtain the product of the homography matrix and the second projection matrix to obtain the first projection matrix.
6. The method according to claim 5, characterized in that The determining, based on the real image and the target virtual image, a homography matrix between the real image and the target virtual image includes: Performing feature extraction on the real image using a feature extraction network to obtain a real feature tensor, and performing feature extraction on the target virtual image using the feature extraction network to obtain a virtual feature tensor; Converting the real feature tensor into a plurality of real feature vectors, and converting the virtual feature tensor into a plurality of virtual feature vectors; Matching the real feature vector and the virtual feature vector using a feature matching network to generate a matching relationship graph, wherein the matching relationship graph is used to represent the corresponding relationship between the real feature vector and the virtual feature vector when the real feature vector and the virtual feature vector are successfully matched; Based on the matching relationship graph, the homography matrix is determined.
7. The method according to claim 6, characterized in that The using a feature matching network to match the real feature vector and the virtual feature vector to obtain a matching relationship graph includes: Matching the real feature vector and the virtual feature vector using a feature matching network to obtain a matching score between the real feature vector and the virtual feature vector; When the matching score is greater than a preset score, it is determined that the real feature vector and the virtual feature vector are matched successfully, and the matching relationship graph is generated based on the real feature vector, the virtual feature vector and the matching score.
8. The method according to claim 6, characterized in that The converting the real feature tensor into a plurality of real feature vectors, and converting the virtual feature tensor into a plurality of virtual feature vectors, comprises: Expanding the real feature tensor according to the spatial dimension to obtain the multiple real feature vectors; The virtual feature tensor is expanded according to the spatial dimension to obtain the multiple virtual feature vectors.
9. The method according to any one of claims 1 to 8, characterized in that The method further comprises: outputting the initial position and the initial posture; In response to an adjustment instruction for adjusting the initial position and the initial posture, determining a target position and a target posture corresponding to the adjustment instruction; Based on the target position and the target posture, the movement of the second image acquisition device is controlled.
10. The method according to any one of claims 1 to 8, characterized in that The determining of a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and device parameters corresponding to the target virtual image includes: outputting the virtual image set and the target virtual image; In response to a selection instruction for selecting other virtual images in the virtual image set, determining a virtual image corresponding to the selection instruction, wherein the other virtual images are virtual images in the virtual image set other than the target virtual image; The first projection matrix is determined based on the real image, the virtual image corresponding to the selection instruction, and device parameters corresponding to the virtual image corresponding to the selection instruction.
11. A calibration method for an image acquisition device, characterized in that: include: Acquire a real image by calling a first interface, wherein the first interface includes a first parameter, a parameter value of the first parameter includes the real image, and the real image is acquired by a first image acquisition device in the real world; Based on the initial position and initial posture of the first image acquisition device, controlling the movement of a second image acquisition device in a virtual world, and obtaining a set of virtual images captured by the second image acquisition device during the movement, wherein the virtual world is a world obtained by simulating the real world; Determining a target virtual image that successfully matches the real image from the virtual image set; determining a first projection matrix of the first image acquisition device based on the real image, the target virtual image, and device parameters corresponding to the target virtual image, wherein the first projection matrix is used to represent a mapping relationship between the virtual world and the real world; The first projection matrix is output by calling a second interface, wherein the second interface includes a second parameter, and a parameter value of the second parameter includes the first projection matrix.
12. An electronic device, characterized in that: include: a memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 11 when running.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the storage medium is located is controlled to execute the method according to any one of claims 1 to 11.
14. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Video fusion method and device based on three-dimensional scene, storage medium and electronic device
CN113870163A
Virtual imaging method and device, nonvolatile storage medium and computer equipment
CN115393497A
Pose initialization method and device of virtual reality head-mounted equipment, equipment and medium
CN117372525A
HMD calibration with direct geometric modeling
EP2966863A1
Virtual image display method and apparatus, electronic device and storage medium
WO2022088918A1