Image fusion processing method and device and electronic equipment

By performing feature extraction and pose calculation in digital twin video fusion, pixel-level fusion of virtual and real scenes is achieved, solving the problem of insufficient detail information in small scene videos and improving the accuracy and realism of the fusion.

CN120852174APending Publication Date: 2025-10-28GUANGZHOU SHIYUAN ELECTRONICS CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410434258.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-11
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing digital twin video fusion technologies cannot capture good detail information in small-scene videos, resulting in poor fusion effects.

Method used

By acquiring images of virtual and real 3D scenes, feature extraction and matching are performed, the pose of the data acquisition device is calculated, and 3D points in the virtual image are projected onto the real image based on the pose and intrinsic parameter information to achieve pixel-level fusion. The color information of the pixel position is replaced to achieve seamless connection.

Benefits of technology

It improves the accuracy and realism of image fusion, ensuring smooth fusion edges, natural details, and providing realistic visual effects, while supporting image fusion in real-time performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852174A_ABST
    Figure CN120852174A_ABST
Patent Text Reader

Abstract

The invention discloses an image fusion processing method and device and electronic equipment. The method comprises the following steps: acquiring a first image corresponding to a preset virtual three-dimensional scene; acquiring a second image in a real three-dimensional scene corresponding to the virtual three-dimensional scene; performing feature extraction and matching based on the first image and the second image to obtain matched feature point pairs; calculating the pose of the data acquisition equipment in the virtual three-dimensional scene according to the matched feature point pairs; projecting a 3D point corresponding to the first image in the virtual three-dimensional scene into a second image according to the pose and the internal reference information of the data acquisition equipment to obtain a target pixel position and target color information corresponding to the 3D point in the second image; and replacing the original color information of the pixel corresponding to the target pixel position in the first image with the target color information to realize fusion of the first image and the second image. The method realizes fusion of virtual and real scenes, and has the characteristics of rich information, accurate position and attitude estimation, real-time performance and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video image processing technology, and in particular to an image fusion processing method, apparatus and electronic device. Background Technology

[0002] In the current development of digital twin technology, digital twins have become an important tool for simulating the digital representation of real-world systems or objects. The goal of digital twin video fusion is to merge multiple digital twin video streams or video clips together to generate a comprehensive, continuous, and realistic digital twin video. The fusion of digital twin video streams or video clips is essentially the fusion of images.

[0003] Existing digital twin video fusion technologies have some limitations in fusing videos from real-world and digital virtual scenes, such as the inability to capture sufficient detail in small-scale video footage. These issues restrict the effectiveness of digital twin video fusion technologies in both real-world and digital virtual scenarios. Summary of the Invention

[0004] The main technical problem addressed by the embodiments of this application is how to improve the accuracy and realism of video fusion.

[0005] To address the aforementioned technical problems, one technical solution adopted in this application is: providing an image fusion processing method, comprising: acquiring a first image corresponding to a preset virtual 3D scene, the first image being configured with original color information; acquiring a second image in a real 3D scene corresponding to the virtual 3D scene, the second image being acquired by a data acquisition device and having a shared viewing area with the first image; performing feature extraction and matching based on the first image and the second image to obtain matched feature point pairs; calculating the pose of the data acquisition device in the virtual 3D scene based on the matched feature point pairs; projecting 3D points corresponding to the first image in the virtual 3D scene onto the second image based on the pose and the intrinsic parameter information of the data acquisition device to obtain the target pixel position and target color information corresponding to the 3D points in the second image; replacing the original color information of the pixels corresponding to the target pixel position in the first image with the target color information to achieve the fusion of the first image and the second image. In this method, by calculating matching feature point pairs and projecting 3D points from the first image onto the second image based on their pose, the pixel positions and color information of the 3D points in the second image can be accurately determined. By replacing the color of the corresponding pixel positions in the first image, pixel-level fusion can be achieved, ensuring a seamless connection between the two images at the pixel level. This pixel-level fusion method provides higher detail preservation and better visual effects. It can synthesize images more accurately, making the fusion edges smoother and the details more natural, thus improving the accuracy and realism of image fusion. Furthermore, the first image in the virtual scene carries the original color information, which can enrich the visual effects of the real scene. By replacing the original color information of the target pixel position with the target color information, the virtual object can be more integrated with the real environment, providing a more realistic visual effect. Moreover, image fusion can be performed in real time, allowing users to observe the performance of virtual objects in the real environment instantly. The virtual-real scene fusion achieved by the method in this embodiment can provide an augmented reality experience and has benefits such as rich information, accurate position and pose estimation, and real-time performance. Furthermore, for video streams, this image fusion processing method can be applied to each frame of the video stream, thereby achieving real-time fusion of the video stream and improving the accuracy and realism of video fusion.

[0006] Optionally, the step of performing feature extraction and matching based on the first image and the second image to obtain matched feature point pairs includes: preprocessing the first image and the second image to obtain preprocessed first image and second image; extracting feature points from the preprocessed first image and second image respectively; for each feature point, calculating its feature descriptor to obtain feature descriptors of the first image and the second image; establishing an index structure for the feature descriptors of the second image; traversing each feature descriptor of the first image and calculating the similarity measure between each feature descriptor in the index structure and the current feature descriptor of the first image; selecting the K most similar feature descriptors to the current feature descriptor of the first image according to the similarity measure, where K is a positive integer; calculating the distance between the current feature descriptor of the first image and each of the K most similar feature descriptors corresponding to the current feature descriptor to obtain the distance between the two feature descriptors; when the distance is less than a preset threshold, determining the feature points corresponding to the two feature descriptors with a distance less than the preset threshold as matched feature point pairs. Establishing a feature descriptor index structure for the second image improves matching efficiency and reduces search time, enabling feature point matching to be completed in a shorter time and improving the algorithm's real-time performance and efficiency. Selecting the K most similar feature descriptors improves matching accuracy, and by setting a preset threshold, only feature point pairs with a distance less than the threshold are selected, further enhancing matching precision and reducing the possibility of false matches. Therefore, this feature extraction and matching method has advantages such as high efficiency and accuracy.

[0007] Optionally, calculating the pose of the data acquisition device in the virtual 3D scene based on the matched feature point pairs includes: converting the pixel coordinates of each feature point in the matched feature point pair into normalized coordinates in the coordinate system of the data acquisition device based on the intrinsic parameter information of the data acquisition device; obtaining the coordinates of each feature point in the matched feature point pair in the virtual 3D scene; selecting a preset number of feature point pairs from the matched feature point pairs, each of the preset number of feature point pairs including the normalized coordinates and the coordinates in the virtual 3D scene; constructing a linear equation for each feature point pair based on the normalized coordinates and the coordinates in the virtual 3D scene corresponding to each feature point pair in the preset number of feature point pairs; forming a system of linear equations from all the linear equations; and calculating the pose of the data acquisition device in the virtual 3D scene based on the system of linear equations. Here, calculating the pose of the data acquisition device in the virtual 3D scene using matched feature points describes the position and orientation of the data acquisition device in the virtual scene, providing an accurate reference for subsequent image fusion and enhancement.

[0008] Optionally, the step of projecting the 3D point corresponding to the first image in the virtual 3D scene onto the second image based on the pose and the intrinsic parameter information of the data acquisition device to obtain the target pixel position and target color information of the 3D point in the second image includes: converting the 3D point corresponding to the first image in the virtual 3D scene into a target 3D point in the coordinate system of the data acquisition device based on the pose and the intrinsic parameter information of the data acquisition device; using the intrinsic parameter information of the data acquisition device, projecting the target 3D point from the coordinate system of the data acquisition device onto the pixel coordinate system of the second image to obtain the target pixel position of the target 3D point in the second image, and obtaining the corresponding target color information based on the target pixel position. Specifically, it can accurately project 3D points of a virtual scene onto an image of a real scene, achieving a precise correspondence between the virtual and real scenes. Through the projection process, the target pixel position of the 3D points in the virtual scene in the second image can be obtained. This provides the pixel coordinates of the virtual object in the real environment, enabling precise alignment and fusion of the virtual object with the real image. In addition, the target color information of the target 3D point in the second image can be obtained, thus obtaining the color information of the virtual object in the real image. This allows for the fusion of the color of the real environment with the virtual object. Moreover, the target color information is determined based on the target pixel position. This pixel-level fusion makes the fusion edges smoother and the details more natural.

[0009] Optionally, converting the 3D points corresponding to the first image in the virtual 3D scene into target 3D points in the coordinate system of the data acquisition device based on the pose and the intrinsic parameter information of the data acquisition device includes: obtaining the pixel points corresponding to the 3D points in the pixel coordinate system of the first image based on the 3D points corresponding to the first image in the virtual 3D scene; and converting each pixel point into a target 3D point in the coordinate system of the data acquisition device based on the pose and the intrinsic parameter information of the data acquisition device. Converting the 3D points in the virtual scene into target 3D points in the coordinate system of the data acquisition device based on the pose and the intrinsic parameter information of the data acquisition device helps to improve the accuracy of localization, pose estimation, and fusion results, thereby providing better image fusion effects.

[0010] Optionally, before performing the step of projecting the 3D points corresponding to the first image in the virtual 3D scene onto the second image based on the pose and the intrinsic parameter information of the data acquisition device, the method further includes: filtering the 3D points corresponding to the first image in the virtual 3D scene to obtain filtered 3D points, which are then used to project onto the second image based on the pose and the intrinsic parameter information of the data acquisition device. Filtering achieves beneficial effects such as data simplification, improved projection accuracy, and improved fusion results, thereby improving the overall algorithm performance and the quality of the fusion results.

[0011] Optionally, the filtering of 3D points corresponding to the first image in the virtual 3D scene includes at least one of the following methods: First method: removing 3D points located behind the data acquisition device after projection onto the data acquisition device's coordinate system; Second method: removing 3D points outside the viewing range of the second image after projection onto the data acquisition device's coordinate system; Third method: for 3D points projected onto the data acquisition device's coordinate system, if multiple 3D points have the same pixel position on the second image, only the 3D point closest to the data acquisition device is retained. By excluding points invisible in the second image, removing 3D points outside the viewing range, and retaining the closest 3D point, the visualization effect, accuracy, and realism of the fusion result can be improved.

[0012] To address the aforementioned technical problems, another technical solution adopted in this application is: providing an image fusion processing apparatus, comprising: a first image acquisition module, configured to acquire a first image corresponding to a preset virtual 3D scene, the first image being configured with original color information; a second image acquisition module, configured to acquire a second image in a real 3D scene corresponding to the virtual 3D scene, the second image being acquired by a data acquisition device and having a shared viewing area with the first image; a feature point determination module, configured to perform feature extraction and matching based on the first image and the second image to obtain matched feature point pairs; a pose calculation module, configured to calculate the pose of the data acquisition device in the virtual 3D scene based on the matched feature point pairs; and a fusion processing module, configured to project the 3D points corresponding to the first image in the virtual 3D scene onto the second image based on the pose and the intrinsic parameter information of the data acquisition device, to obtain the target pixel position and target color information corresponding to the 3D points in the second image; and to replace the original color information of the pixels corresponding to the target pixel position in the first image with the target color information to achieve the fusion of the first image and the second image. In this method, by calculating matching feature point pairs and projecting 3D points from the first image onto the second image based on their pose, the pixel positions and color information of the 3D points in the second image can be accurately determined. By replacing the color of the corresponding pixel positions in the first image, pixel-level fusion can be achieved, ensuring seamless connection between the two images at the pixel level. This pixel-level fusion method provides higher detail preservation and better visual effects. It can synthesize images more accurately, making the fusion edges smoother and the details more natural, thus improving the accuracy and realism of image fusion. Furthermore, the first image in the virtual scene carries the original color information, which can enrich the visual effects of the real scene. By replacing the original color information of the target pixel position with the target color information, the virtual object can be more integrated with the real environment, providing a more realistic visual effect. Moreover, image fusion can be performed in real time, allowing users to observe the performance of virtual objects in the real environment instantly. The device in this embodiment of the application enables the fusion of virtual and real scenes, providing an augmented reality experience with benefits such as rich information, accurate position and pose estimation, and real-time performance. Furthermore, for video streams, this image fusion processing method can be applied to each frame of the video stream, thereby achieving real-time fusion of the video stream and improving the accuracy and realism of video fusion.

[0013] To address the aforementioned technical problems, another technical solution adopted in this application is to provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the image fusion processing method described above. This electronic device has the same beneficial effects as the image fusion processing method described above.

[0014] To address the aforementioned technical problems, another technical solution adopted in this application is to provide a non-volatile computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by an electronic device, the electronic device performs the image fusion processing method described above. This non-volatile computer-readable storage medium has the same beneficial effects as the image fusion processing method described above.

[0015] To address the aforementioned technical problems, another technical solution adopted in this application is to provide a computer program product, comprising a computer program stored on a non-volatile computer-readable storage medium, the computer program including program instructions, which, when executed by an electronic device, cause the electronic device to perform the image fusion processing method described above. This computer program product has the same beneficial effects as the image fusion processing method described above. Attached Figure Description

[0016] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0017] Figure 1 This is a flowchart of a digital twin video fusion method provided in an embodiment of this application;

[0018] Figure 2 This is a flowchart of a method for obtaining matched feature points by performing feature extraction and matching based on the first image and the second image, as provided in an embodiment of this application.

[0019] Figure 3 This is a flowchart of a method for calculating the pose of the data acquisition device in a virtual 3D scene based on the matched feature points, provided in an embodiment of this application.

[0020] Figure 4 This is a schematic diagram of the structure of an image fusion processing device provided in an embodiment of this application;

[0021] Figure 5 This is a schematic diagram of the hardware structure of an electronic device for performing an image fusion processing method provided in an embodiment of this application. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. It should be noted that, unless otherwise specified, the various features in the embodiments of this application can be combined with each other, all within the scope of protection of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different module division or in a different order than that shown in the device schematic diagram or the flowchart. Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application.

[0023] Digital twin technology refers to the concept of simulating, monitoring, and optimizing real-world entities, processes, or systems in real time using digital models. A digital twin combines a physical entity with its digital representation, using technologies such as sensors, data acquisition, and real-time analysis to achieve virtual simulation and monitoring of the physical entity. A digital twin consists of two main components: the physical entity and the digital model. The physical entity can be a product, equipment, building, factory, or entire ecosystem, while the digital model is a virtual representation of these entities. The digital model models the characteristics, attributes, and behaviors of the physical entity and keeps it synchronized with the actual physical entity.

[0024] Currently, various digital twin technology solutions exist, among which some common ones include: physical model-based digital twins, data-driven digital twins, virtual reality-based digital twins, cloud platform-based digital twins, and edge computing-based digital twins. Digital twin technology can be applied to multiple fields and directions, including industrial manufacturing, urban planning and management, energy and environment, healthcare, logistics and supply chain management, transportation, agriculture, architectural design and maintenance, and many others. With continuous technological development and innovation, the application areas of digital twin technology will continue to expand and deepen.

[0025] Video fusion technology, which combines real-world and digital virtual scenes, is a key area of ​​research in digital twins. Specifically, it involves integrating and fusing digital twin models with real-world video data to generate more realistic and accurate simulation and analysis results. By aligning and fusing video data with digital twin models, virtual objects or scenes can be visually interacted with and integrated into the real world within the video. This provides more intuitive and visual simulation and analysis results, allowing for a more accurate evaluation of the model's behavior and performance in a real-world environment.

[0026] In the process of implementing digital twin video fusion technology, a key step is video fusion, which involves merging video information with the twin model. The most intuitive way to achieve this is to merge the rendered view of the twin model with the video footage. Mainstream video fusion techniques are generally suitable for large scenes and wide-angle shots because these scenes have more environmental information and feature points, making camera pose estimation and visual alignment easier. However, in small-scene videos, due to the limited environmental information and feature points, camera pose estimation and visual alignment can become more difficult. This may result in inaccurate alignment between the rendered view of the digital twin model and the real-world video footage, thus affecting the fusion effect and the representation of detailed information.

[0027] Therefore, this application provides an image fusion processing method and apparatus, which is not only applicable to large-scene and wide-angle video images, but also pays more attention to detail in the fusion processing of small-scene video images, making the alignment between the rendered view of the digital twin model and the real-world video image more accurate, and obtaining accurate and realistic fusion results. By calculating the matching feature point pairs and projecting the 3D points in the first image into the second image according to the pose, the pixel position and color information of the 3D points in the second image can be accurately determined; by replacing the color of the corresponding pixel position in the first image, pixel-level fusion can be achieved, ensuring that the two images are seamlessly connected at the pixel level. This pixel-level fusion method can provide higher detail preservation and better visual effects. It can synthesize images more accurately, making the fusion edges smoother and the details more natural, improving the accuracy and realism of image fusion. In addition, the first image in the virtual scene carries the original color information, which can enrich the visual effects of the real scene. By replacing the original color information of the target pixel position with the target color information, the virtual object can be more integrated with the real environment, providing a more realistic visual effect. Furthermore, image fusion can be performed in real time, allowing users to observe the behavior of virtual objects in a real environment instantly. For video streams, this image fusion processing method can be applied to each frame of the video stream, thereby achieving real-time fusion effects and improving the accuracy and realism of video fusion.

[0028] The image fusion processing method provided in this application is applicable to many scenarios, such as virtual reality (VR) and augmented reality (AR) applications, video game development, video production and special effects, virtual navigation, and visualization. In these application scenarios, the hardware devices that may be involved include: data acquisition devices, used to capture images or real-time video of a scene from a set perspective; these can be ordinary cameras or camera equipment specifically designed for VR / AR applications. Computing devices, used for computational tasks such as image processing, visual algorithms, and model inference; these can be computers, embedded systems, cloud servers, etc. Input devices, used for interacting with the system; for example, touchscreens, mice, or gamepads can be used in video production and special effects scenarios. Output devices, used to display the synthesized video results or virtual scenes; these can be computer monitors, television screens, projectors, etc. In some scenarios, audio devices may also be included, such as adding ambient sound effects or user interaction sounds in virtual home decoration. The computing devices can communicate with cameras, input devices, output devices, audio devices, etc.

[0029] In practical applications, appropriate hardware devices can be selected and configured as needed to support the implementation of image fusion processing schemes. The image fusion processing method described in the following embodiments can specifically be executed by the computing device.

[0030] Please see Figure 1 , Figure 1 This is a flowchart of a digital twin video fusion method provided in an embodiment of this application. The method includes:

[0031] S11. Obtain the first image corresponding to the preset virtual three-dimensional scene, wherein the first image is configured with original color information.

[0032] The first image is a digital image of a virtual 3D scene. Acquiring this first image can be achieved through: creating the virtual 3D scene using computer graphics software or libraries; determining the camera's position, pose, and parameters (such as field of view and angle of view), where this step involves camera settings to determine the viewing angle of the scene; next, setting up the rendering, i.e., selecting the rendering algorithm and parameters to define the rendering effect; then running the rendering program to render the virtual 3D scene based on the aforementioned camera and rendering settings, generating an image with color information; finally, extracting the first image corresponding to the preset virtual 3D scene from the rendering result, which includes extracting color information from the rendering result to obtain the image's original color information. Original color information refers to the true color value of each pixel in the image obtained during the rendering process.

[0033] Other methods can also be used to obtain the first image, such as using a real-time graphics rendering engine to render and acquire the image in real time in a virtual scene.

[0034] The first image can have a wide viewpoint, allowing it to include more scene information. Alternatively, it can have a narrow viewpoint, including less scene content. In image fusion processing, an appropriate viewpoint can be selected as needed, providing different observation perspectives to suit different scenarios.

[0035] S12. Obtain a second image from the real three-dimensional scene corresponding to the virtual three-dimensional scene. The second image is acquired by a data acquisition device and has a shared viewing area with the first image.

[0036] Data acquisition equipment refers to devices used to capture video images or video sequences in the real world. These can be cameras, webcams, drone payloads, mobile devices, etc., depending on the application scenario and requirements. Data acquisition equipment uses sensors to capture optical images or video and converts them into digital signals for processing and storage.

[0037] The second image refers to the image captured by the data acquisition device in a real-world scene. This second image is recorded by the camera's optical sensor and reflects the real environment from the camera's position and perspective. The content of the second image can be any object, environment, or scene in a real-world setting; it can be indoor or outdoor, including but not limited to natural scenery, cityscapes, buildings, people, animals, furniture, vehicles, etc. The second image is acquired by the data acquisition device; specifically, it can be a single frame captured by the data acquisition device, or a frame or several frames from a video recorded by the data acquisition device. Besides being captured by the data acquisition device, the second image can also be received and acquired by the data acquisition device, etc.

[0038] The first and second images share a common viewing region, meaning they visually overlap, with certain portions of both images observed from the same viewpoint or perspective. In this embodiment, the existence of this common viewing region enables a seamless transition when fusing the virtual 3D scene and the real image. The pixel-level fusion operation performed within this region ensures a smooth and consistent visual result, avoiding obvious boundaries or discontinuities. Furthermore, the common viewing region facilitates feature extraction and matching between the first and second images, providing crucial foundational information for subsequent image processing tasks.

[0039] S13. Based on the second image and the first image, perform feature extraction and matching to obtain matching feature point pairs.

[0040] Feature extraction and matching algorithms can be used to match feature points in the second image and the first image. The matched feature point pairs are then used to calculate the pose of the data acquisition device in the virtual 3D scene.

[0041] Please refer to Figure 2 Based on the second image and the first image, feature extraction and matching are performed to obtain matched feature point pairs, including:

[0042] S131. Preprocess the first image and the second image to obtain the preprocessed first image and the second image.

[0043] To improve the effectiveness of subsequent feature extraction and matching, the first and second images can be preprocessed first. Preprocessing may include operations such as image denoising, resizing, grayscale conversion, and contrast adjustment.

[0044] S132. Extract feature points from the preprocessed first image and second image respectively.

[0045] Feature points are local regions in an image that have significant properties, such as corners, spots, edges, and line segments.

[0046] In this embodiment, feature points are extracted based on the ORB (Oriented FAST and Rotated BRIEF) algorithm, which is a feature point extraction algorithm that combines FAST corner detection and BRIEF descriptor generation. Specifically, the FAST (Features from Accelerated Segment Test) algorithm can be used to detect corners in the image. The FAST algorithm determines whether a pixel is a corner by comparing the difference in grayscale values ​​between the pixel and its surrounding neighboring pixels, thereby obtaining a set of candidate feature points.

[0047] Other methods can also be used to extract feature points. For example, when the feature points include blobs, blob extraction algorithms include SIFT (Scale Invariant Feature Transform) and DoG (Difference of Gaussians) methods. When the feature points include edges, image segmentation and edge detection can be used to extract edge features. Different types of feature points are suitable for different application scenarios and tasks, and the appropriate feature point type and corresponding extraction method can be selected according to specific needs.

[0048] S133. For each feature point, calculate its feature descriptor to obtain the feature descriptor of the first image and the feature descriptor of the second image.

[0049] A feature descriptor is a numerical vector used to characterize an image or a specific region within an image. The goal of a feature descriptor is to capture the local features and structural information of an image or a region of an image and represent it as a vector with good discriminative power and robustness.

[0050] After obtaining the feature points according to step S132 above, their feature descriptors can be generated using the BRIEF algorithm. The BRIEF descriptor is a binary code that generates a fixed-length binary feature vector by comparing pixel pairs around the feature point. The BRIEF descriptor is characterized by high computational efficiency and good matching robustness. Optionally, to make the ORB feature descriptor rotation-invariant, the orientation of the feature points can be corrected based on the image gradient information around them, thereby aligning the orientations of all feature descriptors to a unified reference direction.

[0051] S134. Establish the index structure of the feature descriptor of the second image.

[0052] The index structure can be a KD tree, a bag-of-words model, a hash table, etc., to improve the efficiency of feature matching, which can accelerate the subsequent similarity measurement and matching process.

[0053] S135. Traverse each feature descriptor of the first image and calculate the similarity measure between each feature descriptor in the index structure and the current feature descriptor of the first image.

[0054] For each feature descriptor in the first image, an index structure is used to search for similar feature descriptors in the second image. Similarity measures can be used, such as Euclidean distance, Hamming distance, or cosine similarity, to calculate the similarity between two feature descriptors.

[0055] S136. Based on the similarity metric, select the K feature descriptors that are most similar to the current feature descriptor of the first image, where K is a positive integer.

[0056] Based on the similarity measurement results, the K feature descriptors that are most similar to the current first image feature descriptor are selected. These most similar feature descriptors are usually considered as potential matching candidates.

[0057] S137. Calculate the distance between the current feature descriptor of the first image and each of the K most similar feature descriptors corresponding to the current feature descriptor to obtain the distance between the two feature descriptors.

[0058] This distance can be calculated using Euclidean distance, Hamming distance, cosine distance, etc., and this application does not specifically limit it in this embodiment.

[0059] S138. When the distance is less than a preset threshold, the feature points corresponding to the two feature descriptors whose distance is less than the preset threshold are determined as a matching feature point pair.

[0060] By applying a distance threshold, matching feature point pairs are identified, thus establishing a correspondence between the first and second images. The preset threshold can be determined through statistical analysis of the dataset. For example, the distance distribution of matching feature point pairs can be calculated, and a suitable threshold can be selected based on statistical indicators such as the mean, standard deviation, or quantiles of the distribution. A smaller preset threshold may lead to more matches but also increases the risk of false matches; conversely, a larger preset threshold may reduce false matches but may also reduce correct matches. Therefore, multiple experiments and adjustments can be conducted to find a suitable preset threshold.

[0061] In this embodiment, establishing a feature descriptor index structure for the second image can improve matching efficiency and reduce search time, thereby enabling feature point matching to be completed in a shorter time and improving the real-time performance and efficiency of the algorithm. Selecting the K most similar feature descriptors can improve matching accuracy, and by setting a preset threshold, only feature point pairs with a distance less than the threshold are selected, further improving matching precision and reducing the possibility of false matches. Therefore, this feature extraction and matching method has advantages such as high efficiency and accuracy.

[0062] S14. Calculate the pose of the data acquisition device in the virtual three-dimensional scene based on the matched feature point pairs.

[0063] The pose is represented by a rotation matrix and a translation vector. The rotation matrix describes the orientation and attitude of the data acquisition device, while the translation vector represents the translation distance of the data acquisition device relative to the virtual 3D scene. Calculating the pose of the data acquisition device in the virtual 3D scene aims to determine its position and orientation within the scene. This is crucial for subsequent image fusion processing, enabling functions such as viewpoint consistency, image correction and calibration, and environmental perception and interaction, thereby improving the quality and effectiveness of image fusion processing.

[0064] Please refer to Figure 3 The step of calculating the pose of the data acquisition device in the virtual 3D scene based on the matched feature point pairs includes:

[0065] S141. Based on the intrinsic parameter information of the data acquisition device, convert the pixel coordinates of each feature point in the feature point pair into normalized coordinates in the coordinate system of the data acquisition device.

[0066] Normalizing the coordinates of feature points makes them unaffected by specific image size and camera intrinsic parameters, facilitating subsequent calculations. The pixel coordinates of a feature point refer to its position coordinates within the corresponding image, which can be represented using an image coordinate system, in pixels. The data acquisition device coordinate system is a coordinate system centered on the data acquisition device, used to describe the device's position and orientation in three-dimensional space, and can consist of three coordinate axes (e.g., X, Y, and Z axes).

[0067] S142. Obtain the coordinates of each feature point in the matched feature point pair in the virtual 3D scene.

[0068] The 3D coordinates corresponding to the 2D points of the data acquisition device can be obtained from the depth map included in the first image, thereby calculating the coordinates of each feature point in the feature point pair in the virtual three-dimensional scene.

[0069] S143. Select a preset number of feature point pairs from the matched feature point pairs, wherein each feature point pair in the preset number of feature point pairs includes the normalized coordinates and the coordinates in the virtual three-dimensional scene.

[0070] For example, four sets of feature point pairs are selected from the matched feature point pairs, and the feature points in each set of feature point pairs include normalized coordinates and coordinates in the virtual three-dimensional scene.

[0071] S144. Based on the normalized coordinates corresponding to each feature point pair in the preset number of feature point pairs and the coordinates in the virtual three-dimensional scene, construct a linear equation for each feature point pair.

[0072] In this system, a set of feature point pairs corresponds to a linear equation, which is established through the correspondence between the feature points.

[0073] S145. Combine all the linear equations into a system of linear equations.

[0074] S146. Calculate the pose of the data acquisition device in the virtual 3D scene based on the linear equations. The pose includes a rotation matrix and a translation vector.

[0075] In this embodiment, the process of obtaining the pose can be considered as solving a Perspective-n-Point (PnP) problem based on the known correspondence between 3D and 2D points. The PnP algorithm is used to solve for the rotation and translation matrices of the data acquisition device. Based on the solved rotation and translation matrices, the pose of the data acquisition device in the virtual 3D scene can be obtained. Here, the rotation matrix represents the rotational attitude of the data acquisition device, and the translation vector represents the translational position of the data acquisition device.

[0076] The PnP algorithm typically requires at least four sets of correspondences between 3D and 2D points, so the preset number can be four. For more accurate pose estimation, more correspondences can be used.

[0077] S15. Based on the pose and the intrinsic parameter information of the data acquisition device, project the 3D point corresponding to the first image in the virtual three-dimensional scene onto the second image to obtain the target pixel position and target color information of the 3D point in the second image.

[0078] Specifically, based on the pose and the intrinsic parameter information of the data acquisition device, the 3D points corresponding to the first image in the virtual 3D scene are converted into target 3D points in the coordinate system of the data acquisition device. Using the intrinsic parameter information of the data acquisition device, the target 3D points are projected from the coordinate system of the data acquisition device onto the pixel coordinate system of the second image to obtain the target pixel position of the target 3D point in the second image, and the corresponding target color information is obtained based on the target pixel position. Specifically, the pixel point corresponding to the 3D point in the virtual 3D scene can be obtained based on the 3D point corresponding to the first image; each pixel point is converted into a target 3D point in the coordinate system of the data acquisition device based on the pose and the intrinsic parameter information of the data acquisition device. In the rendered image, each pixel point typically corresponds to a 3D point in the virtual scene, and this correspondence can be obtained by the rendering engine based on the geometric information of the virtual scene. Pixel points use two-dimensional coordinates to represent their position on the image plane. 3D points in the virtual scene have three-dimensional coordinates, and the pixels in the rendered image are based on the projection results of these 3D points; therefore, they are positions on a two-dimensional plane.

[0079] Specifically, for each 3D point P_r in the first image, a projection transformation is performed using the pose T_c_r and the intrinsic parameter K of the data acquisition device; then, the 3D point P_r is projected onto the second image plane using the intrinsic parameter K of the data acquisition device to obtain the 2D pixel position p_c in the second image. This pixel position represents the projection position of the 3D point on the second image; next, based on the pixel coordinates of p_c, the color information of the corresponding pixel position is obtained in the second image.

[0080] Where p_c = K × T_c_r × P_r;

[0081] p_c represents the 2D pixel position of the second image, P_r represents the 3D coordinates of each 2D point in the first image, K represents the intrinsic parameter information of the data acquisition device, and T_c_r represents the pose of the data acquisition device relative to the first image.

[0082] Through the above steps, based on the camera pose and intrinsic parameters, 3D points in the first image can be projected from the first image to the second image through pose transformation and projection operations. The pixel positions and color information of these 3D points in the second image are then obtained. This enables video fusion, allowing the content of the first image to be superimposed or composited with the second image. Specifically, the above involves using matrix multiplication for 3D point projection, which directly calculates the pixel positions and RGB color information of the points in the second image without traversing all pixels, thus saving computation time and improving the algorithm's efficiency.

[0083] In some embodiments, the method further includes obtaining the intrinsic parameter information of the data acquisition device. The Zhang Zhengyou calibration algorithm can be used to calculate the intrinsic and extrinsic parameters of the data acquisition device. This process includes preparing a calibration board, capturing images of the calibration board, detecting corner points, calculating the intrinsic parameters of the data acquisition device, optimizing the distortion parameters of the data acquisition device, calculating the extrinsic parameters of the data acquisition device, and optimizing the extrinsic parameters of the data acquisition device. For example, if the data acquisition device is a camera, the camera's intrinsic parameters include focal length, principal point, pixel size, lens distortion parameters, etc.; the camera's extrinsic parameters include the camera's pose information, including rotation matrices and translation vectors, etc.

[0084] S16. Replace the original color information of the pixel corresponding to the target pixel position in the first image with the target color information to achieve the fusion of the first image and the second image.

[0085] The above steps obtained the pixel position and color information of 3D points in the first image and their corresponding pixel positions in the second image. This information describes where a pixel in the first image should be located in the second image and what color it should have. Based on this pixel position and color information, the pixel color corresponding to the corresponding pixel position in the first image is replaced with the pixel color of the corresponding position in the second image, thereby achieving the fusion of the second image and the first image.

[0086] This embodiment can accurately project 3D points of a virtual scene onto an image of a real scene, achieving a precise correspondence between the virtual and real scenes. Through the projection process, the target pixel position of the 3D point in the virtual scene in the second image can be obtained. This provides the pixel coordinates of the virtual object in the real environment, enabling precise alignment and fusion of the virtual object with the real image. In addition, the target color information of the target 3D point in the second image is obtained, which allows the color information of the virtual object in the real image to be obtained, thereby achieving the fusion of the color of the real environment with the virtual object. Moreover, the target color information is determined based on the target pixel position. This pixel-level fusion makes the fusion edges smoother and the details more natural.

[0087] The image fusion processing method of this application embodiment can achieve the fusion of real images and images corresponding to virtual 3D scenes. For video fusion of digital twins, each frame in the video stream can be regarded as an independent image, and the image fusion processing method of this application embodiment can be applied to each frame. By fusing real images with virtual 3D scene images, virtual objects or scenes can be mixed with the real world, thereby achieving realistic augmented reality effects or other visual effects in the video stream.

[0088] In some embodiments, before performing step S15 above, the method further includes:

[0089] The 3D points corresponding to the first image in the virtual 3D scene are filtered to obtain filtered 3D points. These filtered 3D points are then projected onto the second image based on the pose and the intrinsic parameters of the data acquisition device. The filtering of the 3D points corresponding to the first image to obtain filtered 3D points includes at least one of the following methods: First method: removing 3D points located behind the data acquisition device after projection onto the data acquisition device's coordinate system; Second method: removing 3D points that are not within the viewing range of the second image after projection onto the data acquisition device's coordinate system; Third method: for 3D points projected onto the data acquisition device's coordinate system, if multiple 3D points have the same pixel position on the second image, only the 3D point closest to the data acquisition device is retained.

[0090] The process involves several steps. First, by projecting the 3D points corresponding to the first image onto the coordinate system of the data acquisition device, it can be determined whether these 3D points are located behind the device. If so, they can be excluded because they will not be visible in the second image. Second, by projecting the 3D points from the first image onto the second image, it can be determined whether these points are within the viewpoint of the second image. If they are not within the viewpoint, they can be excluded because they will not affect the final result during the fusion process. Finally, selecting the 3D points closest to the data acquisition device ensures a more accurate fusion result.

[0091] By filtering the 3D points corresponding to the first image, unwanted or irrelevant points can be removed, improving the effect and accuracy of the fusion process.

[0092] In some embodiments, for each pixel location, a fusion weight can be introduced to control the contribution of the first and second images to the final fusion result. More fine-grained control of the fusion effect can be achieved by adjusting the weights, such as gradual fade-in / fade-out or strengthening the influence of a particular image. Specifically, for example, a weight image of the same size as the input image can be created, where each pixel location corresponds to a weight value. Initially, the weight values ​​for all pixel locations can be set to be equal, indicating that the two images contribute equally to the fusion result. Then, depending on the requirements, the weight values ​​in the weight image can be adjusted using different methods. For example, a gradient function can be used to achieve a gradual fade-in / fade-out effect, causing a particular image to gradually strengthen or weaken at the fusion boundary; alternatively, user interaction or prior knowledge can be used to manually adjust the weight values ​​to strengthen or weaken the influence of a particular image. After adjusting the weights, it is ensured that the weight values ​​in the weight image are within the range of 0 to 1. This can be achieved by normalizing the weight values ​​so that the sum of all weight values ​​equals 1, thus ensuring that the brightness and contrast of the fusion result are not affected by the weight adjustment. Finally, the original input image and the normalized weight image are multiplied pixel by pixel to obtain the weighted image. Specifically, for each pixel location, the pixel values ​​of the first and second images are multiplied by the corresponding weight value, and then the two are added together to obtain the final fusion result.

[0093] By adjusting the weight values ​​of the weighted images, more fine-grained control over the fusion result can be achieved. For example, if you want the second image to dominate the fusion, you can increase the weight value of the corresponding position in the second image; if you want the first image to stand out more in the fusion, you can increase the weight value of the corresponding position in the first image. This allows you to achieve different fusion effects according to your needs, making a certain image more prominent in specific regions or boundaries. This can also enhance the visibility of details, preserve structure and shape, and guide the observer's attention, which is of great significance for the purpose, visual effect, and information delivery of image fusion.

[0094] Please see Figure 4 , Figure 4 This is a schematic diagram of the structure of an image fusion processing device provided in an embodiment of this application. Figure 4 As shown, the image fusion processing device 20 includes: a first image acquisition module 201, a second image acquisition module 202, a feature point determination module 203, a pose calculation module 204, and a fusion processing module 205.

[0095] The first image acquisition module 201 is used to acquire a first image corresponding to a preset virtual 3D scene, the first image being configured with original color information; the second image acquisition module 202 is used to acquire a second image in a real 3D scene corresponding to the virtual 3D scene, the second image being acquired by a data acquisition device and having a shared viewing area with the first image; the feature point determination module 203 is used to perform feature extraction and matching based on the second image and the first image to obtain matched feature point pairs; the pose calculation module 204 is used to calculate the pose of the data acquisition device in the virtual 3D scene based on the matched feature point pairs; the fusion processing module 205 is used to project the 3D points corresponding to the first image in the virtual 3D scene onto the second image based on the pose and the intrinsic parameter information of the data acquisition device, to obtain the target pixel position and target color information corresponding to the 3D points in the second image; and to replace the original color information of the pixels corresponding to the target pixel position in the first image with the target color information to achieve the fusion of the first image and the second image.

[0096] It should be noted that the above-described image fusion processing apparatus can execute the image fusion processing method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in the embodiments of the image fusion processing apparatus can be found in the image fusion processing method provided in the embodiments of this application.

[0097] Please see Figure 5 , Figure 5 This is a schematic diagram of the hardware structure of the electronic device 30 for performing the image fusion processing method provided in the embodiments of this application, as shown below. Figure 5 As shown, the electronic device 30 includes:

[0098] One or more processors 301 and memory 302, Figure 5 Taking a processor 301 as an example, the processor 301 and the memory 302 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.

[0099] The memory 302, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the image fusion processing method in the embodiments of this application. The processor 301 executes various functional applications and data processing of the electronic device by running the non-volatile software programs, instructions, and modules stored in the memory 302, thereby implementing the image fusion processing method of the above-described method embodiments.

[0100] The memory 302 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function. The data storage area may store data created based on the use of the image fusion processing device. Furthermore, the memory 302 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.

[0101] The one or more modules are stored in the memory 302, and when executed by the one or more processors 301, they perform the image fusion processing method in any of the above method embodiments.

[0102] The above-described product can perform the methods provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects for performing the methods. Technical details not described in detail in this embodiment can be found in the methods provided in the embodiments of this application.

[0103] The electronic devices in this application can exist in various forms, including but not limited to: ultra-mobile personal computer devices, mobile communication devices, servers, and other electronic devices with data interaction functions.

[0104] This application provides a non-volatile computer-readable storage medium storing computer-executable instructions that are executed by one or more processors, for example... Figure 5 One of the processors 301 can enable the one or more processors to execute the image fusion processing method in any of the above method embodiments.

[0105] This application provides a computer program product, which includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions that, when executed by the electronic device, enable the electronic device to perform the image fusion processing method described in any of the above method embodiments.

[0106] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0107] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software and a general-purpose hardware platform, or of course, using hardware. Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of this application as described above, which are not provided in detail for the sake of brevity; although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. An image fusion processing method, characterized in that, include: Obtain a first image corresponding to a preset virtual 3D scene, wherein the first image is configured with original color information; Acquire a second image from the real 3D scene corresponding to the virtual 3D scene. The second image is acquired by a data acquisition device and has a shared viewing area with the first image. Based on the first image and the second image, feature extraction and matching are performed to obtain matching feature point pairs; Calculate the pose of the data acquisition device in the virtual 3D scene based on the matched feature point pairs; Based on the pose and the intrinsic parameter information of the data acquisition device, the 3D points corresponding to the first image in the virtual 3D scene are projected onto the second image to obtain the target pixel position and target color information of the 3D points in the second image; The original color information of the pixel corresponding to the target pixel position in the first image is replaced with the target color information to achieve the fusion of the first image and the second image.

2. The method according to claim 1, characterized in that, The step of extracting and matching features based on the first image and the second image to obtain matched feature point pairs includes: The first image and the second image are preprocessed to obtain the preprocessed first image and the second image. Feature points are extracted from the preprocessed first and second images, respectively. For each feature point, its feature descriptor is calculated to obtain the feature descriptor of the first image and the feature descriptor of the second image; Establish the index structure for the feature descriptors of the second image; Traverse each feature descriptor of the first image and calculate the similarity measure between each feature descriptor in the index structure and the current feature descriptor of the first image; Based on the similarity metric, select the K feature descriptors that are most similar to the current feature descriptor of the first image, where K is a positive integer; The distance between the current feature descriptor of the first image and each of the K most similar feature descriptors corresponding to the current feature descriptor is calculated to obtain the distance between the two feature descriptors. When the distance is less than a preset threshold, the feature points corresponding to the two feature descriptors whose distance is less than the preset threshold are determined as a matching feature point pair.

3. The method according to claim 1, characterized in that, The step of calculating the pose of the data acquisition device in the virtual 3D scene based on the matched feature point pairs includes: Based on the intrinsic parameter information of the data acquisition device, the pixel coordinates of each feature point in the matched feature point pair are converted into normalized coordinates in the coordinate system of the data acquisition device. Obtain the coordinates of each feature point in the matched feature point pair in the virtual 3D scene; A preset number of feature point pairs are selected from the matched feature point pairs, and each feature point pair in the preset number of feature point pairs includes the normalized coordinates and the coordinates in the virtual three-dimensional scene; Based on the normalized coordinates corresponding to each feature point pair in the preset number of feature point pairs and the coordinates in the virtual 3D scene, a linear equation is constructed for each feature point pair. Form a system of linear equations from all the aforementioned linear equations; The pose of the data acquisition device in the virtual 3D scene is calculated based on the system of linear equations.

4. The method according to any one of claims 1 to 3, characterized in that, The step of projecting the 3D points corresponding to the first image in the virtual 3D scene onto the second image based on the pose and the intrinsic parameter information of the data acquisition device, to obtain the target pixel position and target color information of the 3D points in the second image, includes: Based on the pose and the intrinsic parameter information of the data acquisition device, the 3D points corresponding to the first image in the virtual 3D scene are converted into target 3D points in the coordinate system of the data acquisition device. Using the intrinsic parameter information of the data acquisition device, the target 3D point is projected from the coordinate system of the data acquisition device onto the pixel coordinate system of the second image to obtain the target pixel position of the target 3D point in the second image, and the corresponding target color information is obtained based on the target pixel position.

5. The method according to claim 4, characterized in that, The step of converting the 3D points corresponding to the first image in the virtual 3D scene into target 3D points in the coordinate system of the data acquisition device based on the pose and the intrinsic parameter information of the data acquisition device includes: Based on the 3D point corresponding to the first image in the virtual 3D scene, obtain the pixel point corresponding to the 3D point in the pixel coordinate system of the first image; Based on the pose and the intrinsic parameter information of the data acquisition device, each pixel is converted into a target 3D point in the coordinate system of the data acquisition device.

6. The method according to claim 4, characterized in that, Before performing the step of projecting the 3D points corresponding to the first image in the virtual 3D scene onto the second image based on the pose and the intrinsic parameter information of the data acquisition device, the method further includes: The 3D points corresponding to the first image in the virtual 3D scene are filtered to obtain filtered 3D points. The filtered 3D points are used to project onto the second image according to the pose and the intrinsic parameter information of the data acquisition device.

7. The method according to claim 6, characterized in that, The step of filtering the 3D points corresponding to the first image in the virtual 3D scene includes at least one of the following methods: The first method: Remove the 3D points located behind the data acquisition device after projection onto the coordinate system of the data acquisition device; The second method: Remove 3D points that are not within the field of view of the second image after being projected onto the coordinate system of the data acquisition device; The third method: For 3D points projected onto the coordinate system of the data acquisition device, if multiple 3D points have the same pixel position on the second image, only the 3D point closest to the data acquisition device is retained.

8. An image fusion processing apparatus, characterized in that, include: The first image acquisition module is used to acquire a first image corresponding to a preset virtual three-dimensional scene, wherein the first image is configured with original color information; The second image acquisition module is used to acquire a second image in the real three-dimensional scene corresponding to the virtual three-dimensional scene. The second image is acquired by a data acquisition device and has a shared viewing area with the first image. The feature point determination module is used to perform feature extraction and matching based on the first image and the second image to obtain matching feature point pairs; The pose calculation module is used to calculate the pose of the data acquisition device in the virtual three-dimensional scene based on the matched feature point pairs. Fusion processing module, used for Based on the pose and the intrinsic parameter information of the data acquisition device, the 3D points corresponding to the first image in the virtual 3D scene are projected onto the second image to obtain the target pixel position and target color information of the 3D points in the second image; The original color information of the pixel corresponding to the target pixel position in the first image is replaced with the target color information to achieve the fusion of the first image and the second image.

9. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the image fusion processing method according to any one of claims 1-7.

10. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-executable instructions, which, when executed by an electronic device, cause the electronic device to perform the image fusion processing method according to any one of claims 1-7.