Image registration model training method, image fusion method and application

By training the image registration model and optimizing the optical flow field with sparse maps, the registration problems caused by differences in view angles between images and inconsistent lighting are solved, and high-quality image fusion is achieved.

CN119963615AActive Publication Date: 2025-05-09SUZHOU INS IMAGE SOFTWARE TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510432563.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-09
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

In the prior art, image registration methods cannot achieve accurate pixel-level alignment, especially when there are large viewing angle differences or light inconsistencies between images, problems such as artifacts, distortions, and loss of details are often encountered.

Method used

By establishing a training sample set, global registration is performed and sparse graphs are generated, the image registration model is trained based on the target cost function, and the sparse graph is used to optimize the current layer optical flow field to achieve multi-layer iterative solution to the matching degree of the optical flow field.

Benefits of technology

Ensure the matching degree of multiple images during registration, reduce the occurrence of abnormal problems such as artifacts, distortions, and details loss, and improve the quality and clarity of image fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963615A_ABST
    Figure CN119963615A_ABST
Patent Text Reader

Abstract

The invention discloses a training method of an image registration model, an image fusion method and application. The image registration model training method comprises the steps that a training sample set is established, and the training sample set comprises image samples obtained by shooting a target object at different visual angles; global registration is carried out on the image sample, and a sparse image is generated; the image registration model is trained based on a target cost function to determine model parameters of the image registration model, the target cost function comprises a reference optical flow field, the image registration model solves the optical flow field based on multi-layer iteration to register image samples, and after a current layer optical flow field is solved, the current layer optical flow field of the current layer is calculated. And optimizing the optical flow field of the current layer by using the sparse graph, and taking the optimized optical flow field as a reference optical flow field of a target cost function when the next layer is iteratively solved. The training method of the image registration model can reduce the abnormal problems of artifacts, distortion, detail loss and the like during registration of a plurality of images due to parallax.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a training method of an image registration model, an image fusion method and applications. Background Art

[0002] In recent years, photographic imaging technology and computational photography have developed rapidly, and multi-camera systems have been widely used in many fields. In the field of consumer electronics, its core functions include widening the field of view, optimizing the depth of field, and integrating cameras with different focal lengths or apertures to achieve computational photography effects such as portrait mode and zoom. In the field of autonomous driving and robotics, multi-camera systems integrate multi-view data to accurately extract depth information, significantly improving the machine's perception of complex environments. In the field of virtual reality (VR) and augmented reality (AR), the system efficiently completes the task of reconstructing 3D objects or scenes by combining multi-view images.

[0003] Image fusion technology can combine multiple images obtained by a multi-camera system into one to present a more realistic visual effect of the subject. When fusing multiple images with different perspectives, it is necessary to align the multiple images to ensure the fusion effect of the multiple images. The existing registration methods cannot achieve precise pixel-level alignment, especially when there are large perspective differences or inconsistent lighting between images, problems such as artifacts, distortion, and loss of details often occur.

[0004] Therefore, in response to the above technical problems, it is necessary to provide an image registration model training method, an image fusion method and an application. Summary of the invention

[0005] The purpose of the present invention is to provide a training method, an image fusion method and application of an image registration model, which can solve the problems of artifacts, distortion, loss of details, etc. that often occur when there are large viewing angle differences or inconsistent lighting between the above-mentioned images.

[0006] In order to achieve the above object, a specific embodiment of the present invention provides a training method for an image registration model, and the technical solution is as follows: A training method for an image registration model, wherein the image registration model is used to register images of different viewing angles, the method comprising: Establishing a training sample set, wherein the training sample set includes image samples taken of a target object at different viewing angles; Performing global registration on the image samples and generating a sparse map, wherein the sparse map includes key feature information in the image samples; The image registration model is trained based on the target cost function to determine the model parameters of the image registration model, wherein the target cost function includes a reference optical flow field. The image registration model is based on multi-layer iterative solution of the optical flow field to align image samples, and after solving the current layer of optical flow field, the sparse graph is used to optimize the current layer of optical flow field as the reference optical flow field of the target cost function when solving the next layer of iterative solution.

[0007] In one or more embodiments of the present invention, optimizing the optical flow field of the current layer by using the sparse graph specifically includes: Generate a guidance field based on the sparse graph, wherein the guidance field includes displacement change information of key features of the sparse graph; A regularization term is constructed based on the guidance field and the current layer optical flow field to perform smoothness constraints on the current layer optical flow field.

[0008] In one or more embodiments of the present invention, the objective cost function includes at least two of a brightness consistency sub-item, a gradient consistency sub-item and a regularization sub-item; the brightness consistency sub-item represents the brightness error between image samples; the gradient consistency sub-item represents the gradient error between image samples; and the regularization sub-item is used to apply regularization to the gradient of the optical flow field.

[0009] In one or more embodiments of the present invention, the image registration model is used to perform local registration on the image samples after global registration.

[0010] In one or more embodiments of the present invention, the optical flow field of the image sample is obtained based on a multi-resolution scale pyramid strategy of optical flow; and / or, The features of the image sample are extracted based on ORB, and the key feature information is screened based on the RANSAC robust matching algorithm to obtain the sparse graph.

[0011] A specific embodiment of the present invention also provides an image fusion method, and the technical solution is as follows: An image fusion method, comprising: Acquire a set of images to be fused, wherein the set of images to be fused includes images of a target object captured at different viewing angles and with different exposures; Fusing the images to be fused, which are images captured at different exposures at the same viewing angle of the target object, to obtain intermediate fused images at each viewing angle; Based on the image registration model obtained by the above training method, the intermediate fused images of each perspective are registered and then fused to obtain a fused image of the target object.

[0012] In one or more embodiments of the present invention, the images to be fused are images captured at different exposures of the target object at the same viewing angle using a weighted average fusion strategy; and the intermediate fused images of each viewing angle are fused using a maximum fusion strategy.

[0013] A specific embodiment of the present invention further provides a training device for an image registration model, and the technical solution is as follows: A training device for an image registration model, wherein the image registration model is used to register images of different viewing angles, the device comprising: An acquisition module is used to establish a training sample set, wherein the training sample set includes image samples taken of a target object at different viewing angles; A first registration module, used for performing global registration on the image samples and generating a sparse map, wherein the sparse map includes key feature information in the image samples; The second registration module is used to train the image registration model based on the target cost function to determine the model parameters of the image registration model, wherein the target cost function includes a reference optical flow field, and the image registration model is based on multi-layer iterative solution of the optical flow field to align image samples, and after solving the current layer of optical flow field, the sparse graph is used to optimize the current layer of optical flow field as the reference optical flow field of the target cost function when solving the next layer of iterative solution.

[0014] A specific embodiment of the present invention further provides an image fusion device, and the technical solution is as follows: An image fusion device, comprising: An acquisition module, used for acquiring a set of images to be fused, wherein the set of images to be fused includes images of a target object captured at different viewing angles and with different exposures; A fusion module, used to fuse the images to be fused, which are images taken of the target object at different exposures at the same viewing angle, to obtain intermediate fused images at each viewing angle; The processing module, based on the image registration model obtained by the training method, registers the intermediate fused images of each perspective and then fuses them to obtain a fused image of the target object.

[0015] A specific embodiment of the present invention further provides an electronic device, and the technical solution is as follows: An electronic device, comprising: at least one processor; and A memory storing instructions, which, when executed by the at least one processor, enables the at least one processor to execute the above-mentioned image registration model training method or the above-mentioned image fusion method.

[0016] A specific embodiment of the present invention further provides a machine-readable storage medium, and the technical solution is as follows: A machine-readable storage medium stores executable instructions, which, when executed, enable the machine to perform the above-mentioned image registration model training method or the above-mentioned image fusion method.

[0017] Compared with the prior art, the present invention has at least one of the following beneficial technical effects: 1. The training method of the image registration model of the present invention trains the image registration model by constructing a target cost function, and at the same time optimizes the optical flow field of the current layer based on a sparse graph, thereby ensuring the matching degree of the optical flow field solved by multiple rounds of iterations, thereby reducing the occurrence of abnormal problems such as artifacts, distortions, and loss of details when multiple images are registered due to parallax.

[0018] 2. The image fusion method of the present invention can perform fusion processing and registration processing on multiple images containing a target object, and can obtain a high-quality 2D image of the target object with rich details. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0020] Figure 1 A diagram showing an implementation environment of a training method for an image registration model and an image fusion method in one embodiment of the present invention; Figure 2 is a flow chart of a method for training an image registration model in one embodiment of the present invention; Figure 3 is a flow chart of an image fusion method in one embodiment of the present invention; Figure 4 A schematic diagram of the layout of a multi-camera system in one embodiment of the present invention; Figure 5 A schematic diagram of the steps of an image fusion method in one embodiment of the present invention; Figure 6 A module diagram of a training device for an image registration model in one embodiment of the present invention; Figure 7 It is a module diagram of an image fusion device in one embodiment of the present invention; Figure 8 It is a hardware structure diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0022] With the rapid development of photographic imaging and computational photography technology, multi-camera imaging systems are becoming increasingly popular in various application fields. A multi-camera system consists of multiple cameras that can collect image information from multiple perspectives to meet diverse application needs. The embodiments of the present invention focus on registering images acquired by multiple cameras to better apply them to subsequent fusion steps, aiming to improve the overall quality of the image, enhance clarity, and achieve a more detailed detail presentation effect.

[0023] With reference Figure 4 and Figure 5 In a specific scene example, a multi-camera system can shoot a target object from different angles, and then perform step-by-step registration operations on the images from different perspectives: ① Global registration of images from different perspectives, ② Pixel registration using the image registration model obtained by the training method of the image registration model of this application. Finally, the registered multi-perspective images are fused to obtain a single high-definition and detailed 2D image.

[0024] Among them, in each embodiment of the present application, the multi-view image used for registration can be directly acquired by a multi-camera system from multiple viewpoints; or, the image at each viewpoint can be obtained by HDR fusion of images of the same viewpoint but different exposures, that is, the images at the same viewpoint are fused into one, and after the fusion of images of the same viewpoint but different exposures is completed, the images are registered.

[0025] Reference Figure 1 , showing a schematic diagram of an implementation environment provided by an exemplary embodiment of the present invention. The implementation environment includes a terminal and a server. The terminal and the server communicate data via a communication network. Optionally, the communication network can be a wired network or a wireless network, and the communication network can be at least one of a local area network, a metropolitan area network, and a wide area network. The image fusion method disclosed in the present invention can be executed by a server, and accordingly, the image registration model can also be deployed in the server.

[0026] Alternatively, the terminal and the server may cooperate to run the image fusion method provided in the embodiment of the present application to complete the fusion of the image. For example, the terminal can be used to perform global registration of images under different viewing angles, or the image registration model can be deployed on the terminal to perform a complete image registration process.

[0027] In a system architecture where the terminal can provide matching computing power, the image fusion method disclosed in the present application can also be directly executed by the terminal, and accordingly, the image registration model can also be deployed on the terminal.

[0028] The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms.

[0029] Reference Figure 2 , introduces a training method for an image registration model in an embodiment of the present invention, wherein the image registration model is used to register images of different perspectives, and the training method for the image registration model includes: S101, establish a training sample set, the training sample set includes image samples of a target object photographed at different viewing angles. Specifically, different cameras may be used to synchronously photograph at the same time to obtain image samples of the target object photographed from different viewing angles.

[0030] S102. Perform global registration on the image samples and generate a sparse graph, wherein the sparse graph includes key feature information in the image samples. In this embodiment, the features of the image samples can be extracted based on ORB (Oriented FAST and Rotated BRIEF), and the key feature information can be filtered based on the RANSAC (Random Sample Consensus) robust matching algorithm to obtain the sparse graph. Among them, ORB is an algorithm for fast feature point extraction and description, and the RANSAC robust matching algorithm is a robust method for estimating mathematical model parameters through random sampling and iteration, which can effectively eliminate the influence of outliers (noise). A sparse graph can be effectively generated based on the ORB algorithm and the RANSAC robust matching algorithm.

[0031] Demonstratively, in ORB feature extraction, the FAST algorithm can be used to detect key points in an image. Compare the grayscale values ​​of a pixel with its surrounding neighboring pixels to quickly find corners; then use the intensity centroid method to calculate the direction of the key point so that the feature is rotationally invariant; then select a set of pixel pairs around the key point, compare their grayscale values, and generate a binary descriptor; finally, match the ORB features extracted from the two images, usually using the Hamming distance as a similarity measure. In RANSAC algorithm matching, a minimum sample set can be randomly selected from the matching point pairs (for example, for homography matrix estimation, at least 4 pairs of matching points are required), the selected sample set can be used to estimate the model parameters, and the errors of all matching point pairs under the current model are calculated, and the point pairs with errors less than the threshold are marked as inliers, and the rest are outliers; repeat the above process multiple times, select the model with the largest number of inliers as the final result, and all inliers can be used to re-estimate the model parameters to further improve the accuracy. In the generation of sparse graphs, the inlier pairs after RANSAC screening can be retained and the outlier pairs can be removed to determine the key points; the screened key point pairs are then used as nodes of the sparse graph and the matching relationships are used as edges.

[0032] S103, training the image registration model based on the target cost function to determine the model parameters of the image registration model, wherein the target cost function includes a reference optical flow field, and the image registration model is based on multi-layer iterative solution of the optical flow field to register image samples, and after solving the current layer of optical flow field, the current layer of optical flow field is optimized using a sparse graph as a reference optical flow field for the target cost function in the next layer of iterative solution. The image registration model is used to perform local registration on the image samples that have been globally registered.

[0033] In this embodiment, the optical flow field of the image sample can be obtained based on the multi-resolution scale pyramid strategy of optical flow. The multi-resolution scale pyramid strategy based on optical flow is a method for optimizing optical flow estimation by constructing an image pyramid (from low resolution to high resolution). It first calculates the rough optical flow on the low-resolution image, and then gradually refines it to the high-resolution layer, thereby improving the calculation efficiency and the accuracy of motion estimation.

[0034] Specifically, the target cost function includes at least two items of brightness consistency sub-item, gradient consistency sub-item and regularization sub-item; the brightness consistency sub-item represents the brightness error between image samples; the gradient consistency sub-item represents the gradient error between image samples; the regularization sub-item is used to regularize the gradient of the optical flow field. It can be understood that the target cost function can include only two items of brightness consistency sub-item and gradient consistency sub-item, or only two items of brightness consistency sub-item and regularization sub-item, or only two items of gradient consistency sub-item and regularization sub-item, or three items of brightness consistency sub-item, gradient consistency sub-item and regularization sub-item.

[0035] The above is the general concept of the target cost function. In this embodiment, the target cost function is further described based on specific parameters, which is not a limitation of the target cost function in this embodiment. The details are as follows: In a specific embodiment, the target cost function includes three items: brightness consistency sub-item, gradient consistency sub-item and regularization sub-item. In this embodiment, the image sample may specifically include a target image and a floating image to be registered with the target image.

[0036]

[0037] in, ω I , ω G and ω R are the weight coefficients of brightness, gradient and regularization respectively; x represents the position coordinate of each pixel in the image; x=(x,y), where x and y are the horizontal and vertical coordinates of the pixel in the image. F(x) represents the optical flow field, specifically the pixel displacement change of the image sample during registration; I r (x) indicates a floating image I r The pixel value at pixel position x. I t (x+F(x)) represents the target image I t In , the pixel value at the pixel position x is shifted by the optical flow field F(x).

[0038] Among them, the sparse graph is used to optimize the current layer optical flow field, specifically including: generating a guiding field based on the sparse graph, wherein the guiding field includes the displacement change information of the key features of the sparse graph; constructing a regularization term based on the guiding field and the current layer optical flow field to perform smoothness constraints on the current layer optical flow field.

[0039]

[0040] in, Indicates thatE (F) The optical flow field F corresponding to the minimum value solution; F guide (x) is the guidance field provided by the sparse graph; γ is the smoothness control parameter; SparseMap represents the sparse graph.

[0041] Among them, the above function can optimize the optical flow field obtained by solving the current layer, and use it as the initial reference optical flow field for solving the optical flow field of the next layer. Specifically, in this embodiment, the optical flow field of the current layer is solved by making the target cost function minimum solution, and the optical flow field value of the current layer is optimized based on the guidance field provided by the sparse graph. Finally, the optimized optical flow field value is used as the initial value of the optical flow field of the next layer iterative solution, so as to gradually complete the precise optimization of the optical flow field and ensure the output result of the model.

[0042] In this embodiment, the constraints of the guiding field provided by the sparse graph can effectively solve the problem caused by the error in the optical flow field estimation in the feature sparse area, thereby improving the robustness of the registration. The sparse area refers to the area in the image sample where the features are sparse and the features are not obvious.

[0043] The training method of the image registration model of this embodiment can effectively reduce abnormal problems such as artifacts and distortions caused by large differences in viewing angles or insufficient local features of the image through the strategies of global registration and local registration, so as to achieve pixel-level precise alignment of the entire image; at the same time, it can also significantly improve the accuracy, versatility and robustness of the registration.

[0044] Reference Figure 3 , an image fusion method in an embodiment of the present invention is introduced. It should be noted in advance that the image registration model involved in the image fusion method of this embodiment can be obtained based on the training method of the above-mentioned image registration model.

[0045] The following specifically introduces an image fusion method in an embodiment of the present invention, which specifically includes: S201. Acquire a set of images to be fused, wherein the set of images to be fused includes images of a target object captured at different viewing angles and with different exposures.

[0046] In this embodiment, a multi-camera system can be used to obtain images of a target object at different viewing angles and exposures. The multi-camera layout can obtain details of the target image at different viewing angles. Figure 4 The number, position, distance and angle of the cameras in this embodiment can be adjusted according to the actual situation to ensure that the detailed information of the target image from different perspectives is presented. The same camera can use multiple exposures (such as low, medium and high) to shoot synchronously at the same time to capture the details of the target object under different lighting conditions.

[0047] S202, fuse the images of the target object captured at different exposures at the same viewing angle in the image set to be fused, and obtain intermediate fused images at each viewing angle. In this embodiment, the images of the same viewing angle with different exposures can be fused by an image fusion model, such as the images can be fused based on a deep learning model. It can be understood that other methods that can be applied to HDR fusion can also be used for the fusion of the image set to be fused in this embodiment.

[0048] With reference Figure 5 , through the multi-camera system, an LDR image sequence with different exposures can be obtained for each viewing angle. The LDR image sequence can be input into the image fusion model to fuse the HDR image, that is, the intermediate fusion image of each viewing angle can be obtained to ensure the quality of the image obtained from each viewing angle. Among them, LDR (Low Dynamic Range) refers to low dynamic range, and its image or video has limited brightness and color range, which is suitable for ordinary display devices. HDR (High Dynamic Range) refers to high dynamic range, which provides a wider brightness and color range and can present more details.

[0049] S203 , based on the image registration model obtained by the training method of the above image registration model, register the intermediate fused images of each perspective and then fuse them to obtain a fused image of the target object.

[0050] With reference Figure 5 , the obtained multi-view intermediate fusion images are registered through the image registration model, and the features in the images at different viewpoints can be aligned to ensure the quality of subsequent image fusion. Among them, in this embodiment, the pixel-level registration of the image can be completed through the cooperation of global registration and local registration, which can solve the problems of distortion and artifacts that occur when the viewpoint difference is large, and performs well when processing three-dimensional objects or images with insufficient local features.

[0051] In this embodiment, the weighted average fusion strategy is adopted in the HDR fusion step of image samples at different exposures, and the maximum fusion strategy is adopted for the fusion of intermediate fusion images of each perspective. The reason is that the HDR fusion of image samples at different exposures needs to obtain images with appropriate exposures, so the weighted average fusion strategy is adopted. The fusion of intermediate fusion images of each perspective needs to integrate the best details, textures and unique information from all perspectives, and finally form a clear and detailed high-quality image, so the maximum fusion strategy is adopted.

[0052] The image fusion method in one embodiment of the present invention overcomes the defect that the prior art cannot effectively utilize the unique detail information of different perspectives in multi-perspective image fusion, and proposes a novel multi-perspective image fusion method. Unlike the traditional method that only focuses on expanding the field of view or extracting depth information, this method maximizes the retention and enhancement of detail information under different perspectives through carefully designed image acquisition, image registration and image fusion strategies, thereby significantly improving the quality and clarity of the final image.

[0053] At the same time, compared with the existing technology, the image fusion method in the present invention further improves the authenticity and consistency of the fused image by coordinating image fusion and image registration, and makes targeted improvements and optimizations for multi-view scenes, especially complex situations such as large parallax.

[0054] Reference Figure 6 , introduces a training device for an image registration model in an embodiment of the present invention. The image registration model is used to register images of different viewing angles, and the device includes an acquisition module 301, a first registration module 302 and a second registration module 303.

[0055] An acquisition module 301 is used to establish a training sample set, which includes image samples taken of a target object at different viewing angles; a first registration module 302 is used to globally register the image samples and generate a sparse map, wherein the sparse map includes key feature information in the image samples; a second registration module 303 is used to train an image registration model based on a target cost function to determine model parameters of the image registration model, wherein the target cost function includes a reference optical flow field, and the image registration model is based on multi-layer iterative solution of the optical flow field to register the image samples, and after solving the current layer of optical flow field, the current layer of optical flow field is optimized using the sparse map as the reference optical flow field of the target cost function when solving the next layer of iterative solution.

[0056] In an optional embodiment, the current layer optical flow field is optimized using a sparse graph, specifically including: generating a guiding field based on the sparse graph, wherein the guiding field includes displacement change information of key features of the sparse graph; constructing a regularization term based on the guiding field and the current layer optical flow field to perform smoothness constraints on the current layer optical flow field.

[0057] In an optional embodiment, the target cost function includes at least two of a brightness consistency sub-item, a gradient consistency sub-item and a regularization sub-item; the brightness consistency sub-item represents the brightness error between image samples; the gradient consistency sub-item represents the gradient error between image samples; and the regularization sub-item is used to apply regularization to the gradient of the optical flow field.

[0058] In an optional embodiment, the image registration model is used to perform local registration on the globally registered image samples.

[0059] In an optional embodiment, an optical flow field of an image sample is obtained based on a multi-resolution scale pyramid strategy of optical flow; and / or, features of the image sample are extracted based on ORB, and key feature information is screened based on a RANSAC robust matching algorithm to obtain a sparse graph.

[0060] Reference Figure 7 , an image fusion device in an embodiment of the present invention is introduced. In this embodiment, the image fusion device includes an acquisition module 401 , a fusion module 402 and a processing module 403 .

[0061] The acquisition module 401 is used to acquire a set of images to be fused, wherein the set of images to be fused includes images of the target object captured at different perspectives and with different exposures; the fusion module 402 is used to fuse the images of the target object captured at the same perspective and with different exposures in the set of images to be fused, to obtain intermediate fused images of each perspective; the processing module 403 is used to align the intermediate fused images of each perspective and fuse them, to obtain a fused image of the target object.

[0062] In an optional embodiment, the processing module 403 further fuses the registered intermediate fused images of each perspective based on the image registration model obtained by the above-mentioned training method.

[0063] In an optional embodiment, the images to be fused are images captured at different exposures of the target object at the same viewing angle using a weighted average fusion strategy; and the intermediate fused images of the viewing angles are fused using a maximum fusion strategy.

[0064] As above Figures 1 to 5 , describes the training method of the image registration model and the image fusion method according to the embodiment of this specification. The details mentioned in the above description of the method embodiment are also applicable to the image fusion device of the embodiment of this specification. The above image fusion device can be implemented by hardware, or by software, or by a combination of hardware and software.

[0065] Reference Figure 8 , shows a hardware structure diagram of an electronic device according to an embodiment of this specification. Figure 8 As shown, the electronic device 50 may include at least one processor 51, a memory 52 (e.g., a non-volatile memory), a memory 53, and a communication interface 54, and the at least one processor 51, the memory 52, the memory 53, and the communication interface 54 are connected together via an internal bus 55. At least one processor 51 executes at least one computer-readable instruction stored or encoded in the memory 52.

[0066] It should be understood that the computer executable instructions stored in the memory 52, when executed, cause at least one processor 51 to perform the above combined operations in various embodiments of the present specification. Figures 1 to 5Describes the various operations and functions.

[0067] In the embodiments of the present specification, the electronic device 50 may include, but is not limited to: a personal computer, a server computer, a workstation, a desktop computer, a laptop computer, a notebook computer, a mobile electronic device, a smart phone, a tablet computer, a cellular phone, a personal digital assistant (PDA), a handheld device, a messaging device, a wearable electronic device, a consumer electronic device, and the like.

[0068] According to one embodiment, a program product such as a machine-readable medium is provided. The machine-readable medium may have instructions (i.e., the above-mentioned elements implemented in the form of software), which, when executed by a machine, causes the machine to perform the above-mentioned combination of various embodiments of this specification. Figure 1-Figure 5 Specifically, a system or device equipped with a readable storage medium may be provided, on which a software program code implementing the functions of any of the above-mentioned embodiments is stored, and a computer or processor of the system or device reads and executes the instructions stored in the readable storage medium.

[0069] In this case, the program code itself read from the machine-readable medium can realize the function of any one of the above embodiments, and thus the machine-readable code and the machine-readable storage medium storing the machine-readable code constitute part of this specification.

[0070] Examples of readable storage media include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code may be downloaded from a server computer or a cloud via a communication network.

[0071] Those skilled in the art should understand that the various embodiments disclosed above can be modified and altered in various ways without departing from the essence of the invention. Therefore, the protection scope of this specification should be defined by the appended claims.

[0072] It should be noted that not all steps and units in the above-mentioned processes and system structure diagrams are necessary, and some steps or units can be ignored according to actual needs. The execution order of each step is not fixed and can be determined as needed. The device structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some units may be implemented by the same physical client, or some units may be implemented by multiple physical clients, or some components in multiple independent devices may be implemented together.

[0073] In the above embodiments, the hardware unit or module can be realized by mechanical or electrical means. For example, a hardware unit, module or processor can include permanent dedicated circuit or logic (such as special processor, FPGA or ASIC) to complete the corresponding operation. The hardware unit or processor can also include programmable logic or circuit (such as general-purpose processor or other programmable processor), which can be temporarily set by software to complete the corresponding operation. Specific implementation (mechanical method or dedicated permanent circuit or temporary circuit) can be determined based on cost and time consideration.

[0074] The specific embodiments described above in conjunction with the accompanying drawings describe exemplary embodiments, but do not represent all embodiments that can be implemented or fall within the scope of protection of the claims. The term "exemplary" used throughout this specification means "used as an example, instance or illustration" and does not mean "preferred" or "having advantages" over other embodiments. For the purpose of providing an understanding of the described technology, the specific embodiments include specific details. However, these technologies can be implemented without these specific details. In some instances, in order to avoid making the concepts of the described embodiments difficult to understand, well-known structures and devices are shown in block diagram form.

[0075] The above description of the present disclosure is provided to enable any person of ordinary skill in the art to implement or use the present disclosure. Various modifications to the present disclosure will be apparent to those of ordinary skill in the art, and the general principles corresponding to the present disclosure may be applied to other variations without departing from the scope of protection of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but is consistent with the widest range of principles and novel features disclosed herein.

[0076] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.

Claims

1. A training method for an image registration model, characterized in that: The image registration model is used to register images of different viewing angles, and the method comprises: Establishing a training sample set, wherein the training sample set includes image samples taken of a target object at different viewing angles; Performing global registration on the image samples and generating a sparse map, wherein the sparse map includes key feature information in the image samples; The image registration model is trained based on the target cost function to determine the model parameters of the image registration model, wherein the target cost function includes a reference optical flow field. The image registration model is based on multi-layer iterative solution of the optical flow field to align image samples, and after solving the current layer of optical flow field, the sparse graph is used to optimize the current layer of optical flow field as the reference optical flow field of the target cost function when solving the next layer of iterative solution.

2. The training method of the image registration model according to claim 1, characterized in that: The sparse graph is used to optimize the optical flow field of the current layer, specifically including: Generate a guidance field based on the sparse graph, wherein the guidance field includes displacement change information of key features of the sparse graph; A regularization term is constructed based on the guidance field and the current layer optical flow field to perform smoothness constraints on the current layer optical flow field.

3. The training method of the image registration model according to claim 1, characterized in that: The target cost function includes at least two items of a brightness consistency sub-item, a gradient consistency sub-item and a regularization sub-item; the brightness consistency sub-item represents the brightness error between image samples; the gradient consistency sub-item represents the gradient error between image samples; and the regularization sub-item is used to apply regularization to the gradient of the optical flow field.

4. The training method of the image registration model according to claim 1, characterized in that: The image registration model is used to perform local registration on the image samples after global registration.

5. The method for training an image registration model according to claim 1, characterized in that: Obtaining the optical flow field of the image sample based on a multi-resolution scale pyramid strategy of optical flow; and / or, The features of the image sample are extracted based on ORB, and the key feature information is screened based on the RANSAC robust matching algorithm to obtain the sparse graph.

6. An image fusion method, characterized in that: include: Acquire a set of images to be fused, wherein the set of images to be fused includes images of a target object captured at different viewing angles and with different exposures; Fusing the images to be fused, which are images captured at different exposures at the same viewing angle of the target object, to obtain intermediate fused images at each viewing angle; Based on the image registration model obtained by the training method according to any one of claims 1 to 5, the intermediate fused images of each perspective are registered and then fused to obtain a fused image of the target object.

7. The image fusion method according to claim 6, characterized in that: The method specifically comprises: Using a weighted average fusion strategy, the images to be fused are images captured at different exposures of the target object at the same viewing angle; and The intermediate fusion images of the various perspectives are fused using a maximum fusion strategy.

8. A training device for an image registration model, characterized in that: The image registration model is used to register images of different viewing angles, and the device comprises: An acquisition module is used to establish a training sample set, wherein the training sample set includes image samples taken of a target object at different viewing angles; A first registration module, used for performing global registration on the image samples and generating a sparse map, wherein the sparse map includes key feature information in the image samples; The second registration module is used to train the image registration model based on the target cost function to determine the model parameters of the image registration model, wherein the target cost function includes a reference optical flow field, and the image registration model is based on multi-layer iterative solution of the optical flow field to align image samples, and after solving the current layer of optical flow field, the sparse graph is used to optimize the current layer of optical flow field as the reference optical flow field of the target cost function when solving the next layer of iterative solution.

9. An image fusion device, characterized in that: include: An acquisition module, used for acquiring a set of images to be fused, wherein the set of images to be fused includes images of a target object captured at different viewing angles and with different exposures; A fusion module, used to fuse the images to be fused, which are images taken of the target object at different exposures at the same viewing angle, to obtain intermediate fused images at each viewing angle; A processing module, based on the image registration model obtained by the training method according to any one of claims 1 to 5, registers the intermediate fused images of each perspective and then fuses them to obtain a fused image of the target object.

10. An electronic device, characterized in that: include: at least one processor; as well as A memory storing instructions, which, when executed by the at least one processor, causes the at least one processor to execute the image registration model training method as described in any one of claims 1 to 5 or the image fusion method as described in any one of claims 6 to 7.

11. A machine-readable storage medium, characterized in that: It stores executable instructions, which, when executed, enable the machine to perform the image registration model training method as described in any one of claims 1 to 5 or the image fusion method as described in any one of claims 6 to 7.

Citation Information

Patent Citations

  • Improved optical flow field model algorithm

    CN108022261A

  • Semantic mapping and positioning method based on priori laser point cloud and depth map fusion

    CN112258618A

  • Coarse-to-fine image dense matching method, system and device and storage medium

    CN114429555A

  • Target detection model training and target detection method and device

    CN117152560A

  • Fluid optical flow velocity measurement method and system for complex illumination change scene

    CN117710416A