Image splicing method and device, electronic equipment, storage medium and program product

By acquiring images through a camera at the end of a robotic arm and calculating global transformation and local homography matrices, the problems of low efficiency and unstable accuracy of manual image stitching are solved, achieving efficient and high-quality image stitching.

CN121235902APending Publication Date: 2025-12-30CHINA MOBILE ZIJIN INNOVATION INST CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511412118.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

In existing technologies, image stitching relies on manual operation, which is inefficient and has unstable accuracy, making it difficult to process large-size or multi-source images.

Method used

Multiple source images are acquired by a camera at the end of a robotic arm, and the global transformation matrix and local homography matrix are determined. Image stitching is then performed by combining pixel-level mapping. By utilizing the precise motion control of the robotic arm and the efficient acquisition of the camera, the viewpoint is kept stable and the overlapping area is controllable, thus achieving precise spatial mapping between images.

Benefits of technology

It improves the accuracy and processing efficiency of image stitching, enhances the structural integrity and local detail clarity of the generated target image, reduces stitching errors, and achieves high-quality image stitching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121235902A_ABST
    Figure CN121235902A_ABST
Patent Text Reader

Abstract

The invention discloses an image splicing method and device, electronic equipment, a storage medium and a program product, and relates to the technical field of image data processing. The method comprises the steps that a plurality of source images of a target object are collected, a global transformation matrix among the source images is determined, then a plurality of image groups are determined, and each image group comprises a first source image and a second source image to be spliced; determining a plurality of target feature point pairs in the first source image and the second source image in each image group, respectively dividing the first source image and the second source image into a plurality of image windows with overlapping areas, and determining a local homography matrix of the image windows according to the target feature point pairs in the image windows and the global transformation matrix; and determining a second pixel point in the second source image according to the local homography matrix associated with the first pixel point in the first source image, and splicing the first source image and the second source image according to the plurality of first pixel points and the corresponding second pixel points to obtain a target image of the target object so as to improve the image splicing effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image data processing technology, and in particular to an image stitching method, apparatus, electronic device, storage medium, and program product. Background Technology

[0002] Images, as an important tool for intuitively recording information and recreating scenes, have been deeply integrated into all aspects of people's lives and production.

[0003] Most related technologies rely on manual image stitching, where operators visually observe the overlapping areas of two or more source images and manually adjust their position, angle, and scaling to align the overlapping areas and create a complete stitched image. This manual stitching method heavily depends on the operator's experience and visual judgment, making it time-consuming and labor-intensive. Furthermore, as the complexity of the scene increases, manual stitching suffers from limitations such as low efficiency, unstable accuracy, and difficulty in handling large-size or multi-source images. Therefore, there is an urgent need for an image stitching method that improves both accuracy and processing efficiency. Summary of the Invention

[0004] This invention provides an image stitching method, apparatus, electronic device, storage medium, and program product to solve the problems of low efficiency and unstable accuracy in manual image stitching in related technologies.

[0005] According to one aspect of the present invention, an image stitching method is provided, the method comprising:

[0006] Multiple source images of the target object are acquired by a camera set at the end of the robotic arm. A global transformation matrix between the multiple source images is determined. Multiple image groups are determined based on the multiple source images. The image groups include a first source image and a second source image to be stitched together. Partial image regions in the first source image and the second source image correspond to the same object region of the target object.

[0007] For each image group, multiple pairs of target feature points in the first source image and the second source image in the image group are determined, and the first source image and the second source image are divided into multiple image windows with overlapping regions, and the local homography matrix corresponding to the image window is determined according to the pairs of target feature points in the image window and the global transformation matrix.

[0008] For a first pixel in the first source image, a second pixel in the second source image is determined based on the local homography matrix associated with the first pixel. The first source image and the second source image are then stitched together based on multiple first pixels and their corresponding second pixels to obtain a target image of the target object.

[0009] According to another aspect of the present invention, an image stitching apparatus is provided, the apparatus comprising:

[0010] The image group determination module is used to acquire multiple source images of a target object through a camera set at the end of a robotic arm, determine a global transformation matrix between the multiple source images, and determine multiple image groups based on the multiple source images. The image group includes a first source image and a second source image to be stitched together, and a portion of the image regions in the first source image and the second source image correspond to the same object region of the target object.

[0011] The window local homography matrix determination module is used to determine multiple target feature point pairs in the first source image and the second source image in each image group, and to divide the first source image and the second source image into multiple image windows with overlapping regions, and to determine the local homography matrix corresponding to the image window based on the target feature point pairs in the image window and the global transformation matrix.

[0012] The target image acquisition module is used to determine a second pixel in the second source image based on the local homography matrix associated with the first pixel in the first source image, and to stitch the first source image and the second source image together based on multiple first pixels and their corresponding second pixels to obtain a target image of the target object.

[0013] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the image stitching method according to any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the image stitching method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the image stitching method as described in any of the embodiments of this disclosure.

[0019] The technical solution of this invention first involves acquiring multiple source images of a target object using a camera mounted at the end of a robotic arm, determining a global transformation matrix between the source images, and defining multiple image groups based on the source images. Each image group includes a first source image and a second source image to be stitched together, with some image regions in the first and second source images corresponding to the same object region of the target object. The precise movement of the robotic arm controls the camera's shooting angle and position, ensuring stable viewing angles and controllable overlapping areas in the acquired source images, laying a high-quality data foundation for subsequent stitching. Determining the global transformation matrix between the multiple source images establishes an overall spatial relationship between the images, reducing the accumulation of errors in subsequent local processing. Next, for each image group, multiple target feature point pairs in the first and second source images within the image group are determined. The first and second source images are then divided into multiple image windows with overlapping areas, and the local homography corresponding to each image window is determined based on the target feature point pairs in the image window and the global transformation matrix. The homography matrix can improve the matching accuracy of the correspondence between images by using multiple target feature points. It divides image windows with overlapping regions, calculates the local homography matrix corresponding to each image window, and accurately adapts to subtle local changes within the window while following global transformation rules, thus improving the accuracy of the transformation matrix in describing the spatial relationship of the image. Finally, for the first pixel in the first source image, the second pixel in the second source image is determined based on the local homography matrix associated with the first pixel. The first source image and the second source image are then stitched together based on multiple first pixels and their corresponding second pixels to obtain the target image of the target object. This achieves pixel-level accurate spatial mapping by stitching images together using multiple sets of spatial correspondences between first pixels and corresponding second pixels. This reduces stitching errors through mutual verification of a large amount of pixel-level data, ensuring the overall structural integrity and local detail clarity of the generated target image. By balancing accuracy and efficiency through a transformation matrix calculation that combines global and local approaches, and achieving high-quality stitching with pixel-level mapping, the target image of the target object can be accurately presented.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of an image stitching method provided according to Embodiment 1 of the present invention;

[0023] Figure 2 This is a flowchart of an image stitching method provided in Embodiment 2 of the present invention;

[0024] Figure 3a This is a flowchart of an image stitching method for use in embodiments of the present invention, according to a third embodiment of the present invention.

[0025] Figure 3b This is a schematic diagram of an image stitching method for use in embodiments of the present invention, according to a third embodiment of the present invention;

[0026] Figure 3c This is a schematic diagram of image detection model processing provided in the image stitching method of the present invention according to Embodiment 3 of the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of an image stitching device according to Embodiment 4 of the present invention;

[0028] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the image stitching method of this invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0033] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0034] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0035] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0036] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0037] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0038] Example 1

[0039] Figure 1 This is a flowchart of an image stitching method provided in Embodiment 1 of the present invention. This embodiment is applicable to stitching images of a target object. The method can be executed by an image stitching device, which can be implemented in hardware and / or software, optionally through an electronic device, such as a mobile terminal, PC, or server. Figure 1 As shown, the method may specifically include:

[0040] S110. Multiple source images of the target object are acquired by a camera set at the end of the robotic arm, a global transformation matrix between the multiple source images is determined, and multiple image groups are determined based on the multiple source images. The image groups include a first source image and a second source image to be stitched together, and some image regions in the first source image and the second source image correspond to the same object region of the target object.

[0041] In this embodiment of the invention, the end effector of the robotic arm can be a carrier on which an actuator such as a camera is mounted. Its position and orientation can be adjusted according to actual business needs to enable the camera to capture images of the target object from different angles and positions. The camera refers to a device installed at the end effector of the robotic arm that can be used to acquire image information of the target object, providing source image data for subsequent image stitching. The target object can be understood as the specific object captured by the camera. The source image can be understood as the original image acquired by the camera for image stitching. Specifically, the source image can be a partial image of the target object. By stitching multiple source images, an image of a larger area of ​​the target object can be stitched together; for example, a global image of the target object can be stitched together. Wherein, at least the source images acquired at adjacent acquisition positions include overlapping areas, and all source images are acquired from the target object. That is, some image content in the two source images to be stitched is the same.

[0042] Optionally, for the target object, the camera's shooting position can be set to ensure that the robotic arm can drive the camera to shoot the target object from multiple perspectives. Then, the robotic arm can be moved according to the shooting position, and multiple source images of the target object can be acquired by the camera set at the end of the robotic arm. The global transformation matrix used to describe the overall spatial position transformation relationship between the multiple source images can be further determined to reflect the relative position and orientation of different source images in the coordinate system from a macroscopic perspective, providing a basic reference for the subsequent determination of the local homography matrix.

[0043] It should be noted that, at least at adjacent acquisition positions, the field of view of the camera at the end of the robotic arm partially overlaps. This means that the source images to be stitched contain some identical image content, allowing image stitching to be achieved by understanding the correspondence between this overlapping content in the two source images.

[0044] Specifically, determining the global transformation matrix between multiple source images includes: determining the global transformation matrix between source images based on the hand-eye transformation matrix; or, when the target object has a known geometric structure, detecting the corresponding geometric features in the source images, establishing a correspondence, and calculating the global transformation matrix by combining the parameters; or, using the phase correlation method, obtaining the translation amount through Fourier transform, and determining the global transformation matrix after obtaining the rotation and scaling parameters through multi-scale calculation, etc., without making specific limitations here.

[0045] The hand-eye transformation matrix refers to a matrix used to establish the transformation relationship between a first coordinate system and a second coordinate system. Taking the hand-eye transformation matrix as an example, based on the above scheme, optionally, determining the global transformation matrix between multiple source images includes: determining the first pose matrix of the robotic arm end effector in the first coordinate system, and determining the second pose matrix of the camera in the second coordinate system based on the first pose matrix and the hand-eye transformation matrix; wherein, the hand-eye transformation matrix is ​​used to indicate the transformation relationship between the first coordinate system and the second coordinate system; and determining the global transformation matrix between multiple source images based on the second pose matrix of the multiple source images.

[0046] The first coordinate system can be understood as the coordinate system used to describe the pose of the robotic arm's end effector. It can be associated with the base of the robotic arm or a fixed component, serving as a reference coordinate system for determining the first pose matrix of the robotic arm's end effector. The first pose matrix is ​​the matrix used to describe the position and attitude of the robotic arm's end effector in the first coordinate system. The first pose matrix can include the coordinate information and attitude angle information of the robotic arm's end effector in space. The second coordinate system is the coordinate system used to describe the camera's pose. It can be constructed using the optical center of the camera or a specific component as the origin. The second pose matrix is ​​the matrix used to describe the camera's position and attitude in the second coordinate system. It can be calculated based on the first pose matrix of the robotic arm's end effector and the hand-eye transformation matrix. By using the second pose matrices corresponding to multiple source images, the pose differences when the camera captures different source images can be reflected, thereby determining the global transformation matrix between multiple source images.

[0047] Specifically, the first pose matrix of the robotic arm end effector in the first coordinate system can be determined as follows: Based on the pose matrix of the robotic arm's end effector Hand-eye conversion matrix Through coordinate transformation, the second pose matrix of the camera in the second coordinate system is calculated. ,Right now The second pose matrix includes the camera's position matrix and pose information.

[0048] From a spatial geometry perspective, the pose (such as rotation and orientation) of an object in three-dimensional space needs to be fully described by three rotational components: rotation around the x-axis, y-axis, and z-axis. A 3×3 matrix structure can hold rotational information in these three dimensions. By combining elements in the matrix, the pose transformation relationship of the camera in three-dimensional space can be accurately expressed. Therefore, a 3×3 submatrix is ​​used to describe the camera's pose.

[0049] Furthermore, the second pose matrix may include a rotation matrix describing the camera's pose information and a translation matrix describing the camera's position information. The second pose matrix can be represented by the following formula:

[0050] ;

[0051] in, This represents the camera's second pose matrix, which is a 4×4 matrix that integrates the camera's pose and position information. It is a 3×3 rotation matrix that represents the camera's pose in three-dimensional space, such as rotation and orientation; This is a 3×1 translation matrix, representing the camera's position in 3D space, specifically the translation along the x, y, and z axes. Below the matrix... The homogeneous coordinate system provides a standardized form, enabling translation and rotation transformations to be calculated within the same mathematical framework, thus ensuring the consistency of matrix operations.

[0052] Optionally, for multiple acquired source images, their corresponding second camera pose matrices can be determined separately. And determine the global transformation matrix between the multiple source images based on the second pose matrix of the multiple source images. .

[0053] Specifically, the global transformation matrix can be determined based on the following formula:

[0054] ;

[0055] in, For the first The source image and the first The global transformation matrix between the source images; For the first The second pose matrix corresponding to each source image; For the first The second pose matrix corresponding to each source image; , This represents the number of source images acquired.

[0056] If the points in source image 1 The corresponding point in source image 2 is ,but It can be represented as Its global transformation matrix is .

[0057] By calculating the second pose matrix of the camera using the first pose matrix of the robotic arm end effector and the hand-eye transformation matrix, the high-precision pose information of the robotic arm can be transformed into a spatial position reference for the camera, effectively reducing the matrix solution error caused by insufficient image features. Based on the camera pose matrix, the global transformation matrix between source images can be determined, and the spatial position correlation of images can be directly established, greatly improving the efficiency and accuracy of global matrix calculation.

[0058] Furthermore, a first source image and a second source image to be stitched can be selected from multiple source images. Partial image regions in the first and second source images correspond to the same object region of the target object, thus determining multiple image groups. By using the commonly captured object region corresponding to the target object in the first and second source images, which overlaps in the two source images, overlap matching can be achieved during image stitching, establishing the correspondence between pixels in the two source images.

[0059] S120. For each image group, determine multiple target feature point pairs in the first source image and the second source image in the image group, and divide the first source image and the second source image into multiple image windows with overlapping regions, and determine the local homography matrix corresponding to the image window based on the target feature point pairs in the image window and the global transformation matrix.

[0060] Here, a target feature point pair refers to a pair of feature points extracted from the first source image and the second source image, which can match each other. Optionally, the two feature points in each feature point pair come from the first source image and the second source image, respectively, and correspond to the same object region on the target object, providing key data for determining the local homography matrix.

[0061] To prevent feature point matching errors, the target feature point pairs can be processed. Optionally, determining multiple target feature point pairs in the first source image and the second source image in the image group includes: performing pixel coordinate correction on the first source image and the second source image in the image group according to the camera intrinsic parameter matrix and camera distortion parameters of the camera; performing brightness correction on the pixel coordinate-corrected first source image and / or second source image to reduce the brightness difference between the first source image and the second source image; and determining multiple feature points in the brightness-corrected first source image and the second source image.

[0062] The camera intrinsic parameter matrix can be understood as a matrix related to the camera's internal parameters. It may include, but is not limited to, parameters such as focal length and principal point coordinates. The parameters of the camera intrinsic parameter matrix are determined by the camera's hardware structure. When performing pixel coordinate correction on the source image, the camera intrinsic parameter matrix is ​​used to eliminate image distortion caused by the camera's internal parameters. Camera distortion parameters can be understood as parameters describing the degree of distortion generated when the camera captures an image. Due to manufacturing errors in the camera's optical system, captured images may exhibit radial distortion, tangential distortion, etc. Camera distortion parameters provide a quantitative description of these distortion phenomena, and can be used to correct distorted images during pixel coordinate correction.

[0063] Specifically, for each image group, a mathematical model can be established based on the camera intrinsic parameter matrix and camera distortion parameters. The coordinates of each pixel in the first source image and the second source image are recalculated respectively. The coordinates of the pixels in the source image are adjusted and corrected. The image distortion caused by the defects of the camera's own optical system is eliminated according to the inverse process of image distortion. The corrected image pixel coordinates more accurately reflect the actual spatial position of the target object, providing more accurate image data for subsequent feature point extraction and image stitching.

[0064] Brightness difference refers to the degree of difference in brightness between the first source image and the second source image. Due to differences in lighting conditions, exposure parameters, and other factors during camera shooting, brightness differences may exist between multiple source images. This can lead to obvious brightness discontinuities in the overlapping areas of the stitched image, affecting the visual effect of the image stitching. Brightness correction can reduce this brightness difference, making the stitched target image more consistent and natural in brightness.

[0065] Specifically, for the first source image and / or the second source image after pixel coordinate correction, the brightness distribution of the first source image and / or the second source image can be analyzed by histogram, and the brightness difference between the first source image and the second source image can be corrected by combining grayscale correction method, and multiple feature points in the first source image and the second source image after brightness correction can be determined.

[0066] By correcting pixel coordinates using the camera intrinsic parameter matrix and distortion parameters, the interference of camera optical distortion on pixel position can be eliminated, ensuring the authenticity and reliability of feature point coordinates. Brightness correction reduces the brightness difference between two source images, avoiding the misidentification of brightness differences as edge features by traditional feature extraction algorithms due to uneven illumination, thus improving the accuracy of feature point extraction. Dual correction preprocessing provides high-quality image data for subsequent feature point pair determination, reduces matching errors caused by preprocessing omissions, and improves the reliability of feature point pairs.

[0067] Optionally, the first source image and the second source image can be divided into multiple image windows with overlapping regions. The following segmentation methods can be used: according to a preset fixed size and step size, image windows are generated by sliding from left to right and from top to bottom on the image plane, where the step size is smaller than the window size to form overlapping regions; or, based on the feature distribution of the target object in the image, the size and position of the window are adaptively determined so that each window contains as complete a local feature region as possible and adjacent windows overlap; or, based on the pixel gradient information of the image, a larger window size is used in areas with gentle gradient changes and a smaller window size is used in areas with drastic gradient changes, while ensuring that there are overlapping regions between the windows, etc., without specific limitations.

[0068] In this context, an image window can be understood as a sub-image block of a certain size with overlapping areas, into which the source image is divided. By processing each image window, the image transformation relationship of local regions can be determined more accurately, improving the precision of image stitching. The overlapping area can be understood as the region where adjacent image windows overlap each other during the image window division process. The overlapping area provides a basis for the transition and fusion between image windows and the stitching matching between source images.

[0069] The size and overlap of the image window are both adjustable parameters. The window size has a significant impact on the effect of capturing local image information and computational efficiency. If a smaller image window is used, the deformation of local areas of the image can be captured more accurately, and the adaptability to changes in local details can be improved. However, the overall computational load will increase due to the increase in the number of windows. If a larger image window is used, the number of windows can be reduced to improve computational efficiency and reduce data processing time. However, if the window coverage is too wide, it may be difficult to accurately reflect the complex deformation and detail changes of local areas, and the accuracy of the characterization of local information will decrease.

[0070] Taking step size as an example, optionally, the window size can be set to... The overlap rate is Then, the step size can be determined based on the window size and overlap rate, and the window can be divided on the image according to the determined step size.

[0071] Optionally, if the number of image windows obtained after partitioning the first source image and the second source image is different, a spatial correspondence between the windows can be established based on the global transformation matrix of the two images, and then the local homography matrix can be calculated. First, the source image with more windows can be used as a reference. Based on the global transformation matrix, the mapping region of each image window in the other source image can be determined. Then, in the other source image, image windows can be partitioned around this mapping region. The size and overlap of the partitioned image windows must be consistent with the existing image windows, so that the number of image windows in the first source image and the second source image match and their spatial positions correspond one-to-one. Alternatively, in the source image with fewer windows, the partitioned image windows can be subdivided. The size of the subdivided sub-windows must be consistent with the window size of the other source image, and the sub-windows must maintain a preset overlap. By increasing the number of windows, a dual match in the number and size of windows in the first and second source images can be achieved, ensuring that the local homography matrix can be accurately calculated for each image window subsequently. No specific limitations are made here.

[0072] The local homography matrix is ​​a matrix used to describe the local spatial transformation relationship between pixels in the first and second source images within each image window. The local homography matrix can be determined based on the global transformation matrix, combined with the target feature point pairs within the image window, and can more accurately reflect the transformation of local regions of the image.

[0073] Based on the above scheme, optionally, determining the local homography matrix corresponding to the image window according to the target feature point pairs in the image window and the global transformation matrix includes: using the global transformation matrix as the initial homography matrix, determining the local homography matrix corresponding to the image window according to the target feature point pairs in the image window and the least squares method.

[0074] The initial homography matrix refers to the matrix used as an initial reference when determining the local homography matrix corresponding to the image window. In this embodiment of the invention, the global transformation matrix can be used as the initial homography matrix to provide a basis for subsequent optimization of the local homography matrix by combining target feature point pairs.

[0075] Specifically, for each image window, when determining its corresponding local homography matrix, the least squares method can be used for calculation. Based on the initial homography matrix as a reference, combined with the target feature point pairs extracted within the image window, an objective function is constructed and the sum of squared errors is minimized. By solving the minimization objective function, the optimal function matching relationship of the data is solved, and finally a local homography matrix that is highly adapted to the local transformation features of the image window is obtained, ensuring that it can accurately reflect the spatial transformation law of the image within the window.

[0076] Continuing with source image 1 and source image 2 as examples, and using the target global transformation matrix as the initial homography matrix, assume that the target feature point pairs within the image window are... , Given the number of target feature point pairs within the image window, a minimum objective function is constructed and solved using the least squares method. ,in, Representing the Euclidean distance, we can obtain the th partitioned value. Local homography matrix of an image window , , The total number of image windows obtained after dividing the source image.

[0077] By using the global transformation matrix as the initial value of the local homography matrix, a reasonable initial range can be set for solving the local matrix based on the overall spatial rules provided by the global matrix. This effectively avoids getting stuck in local optima during least squares iteration, reduces the number of computation iterations, and improves solution efficiency. At the same time, by combining the target feature point pairs within the window with the least squares method, the local subtle transformations within the window can be accurately adapted through data fitting. This ensures that the local homography matrix not only conforms to the overall spatial correlation of the image but also accurately portrays local differences, significantly improving the accuracy of local transformation description and providing a reliable basis for subsequent pixel-level mapping.

[0078] S130. For a first pixel in the first source image, determine a second pixel in the second source image based on the local homography matrix associated with the first pixel, and stitch the first source image and the second source image together based on multiple first pixels and their corresponding second pixels to obtain a target image of the target object.

[0079] Here, the first pixel refers to any pixel in the first source image used for stitching and matching. Through its associated local homography matrix, the corresponding second pixel can be found in the second source image, thus achieving pixel-level stitching of the two source images. The second pixel refers to the pixel in the second source image that corresponds to the first pixel, determined by the first pixel and its associated local homography matrix. The target image refers to the image that represents the target object, obtained by stitching and fusing the first and second source images according to their pixel correspondence.

[0080] Based on the above scheme, optionally, determining the second pixel in the second source image according to the local homography matrix associated with the first pixel includes: determining the target homography matrix corresponding to the first pixel according to the local homography matrix corresponding to the image window where the first pixel is located, and determining the second pixel in the second source image according to the first pixel and the target homography matrix.

[0081] The target homography matrix refers to the homography matrix used to determine the second pixel corresponding to the first pixel in the second source image. The target homography matrix of the first pixel can be determined based on the local homography matrix corresponding to the image window where the first pixel is located, and can accurately reflect the local transformation relationship of the first pixel.

[0082] Because image windows have overlapping areas, a single pixel may simultaneously fall within the range of multiple image windows. Optionally, the target homography matrix of the first pixel can be obtained by weighted fusion of the local homography matrices of the multiple windows in which it resides.

[0083] After calculating the local homography matrix of each image window, the target homography matrix of each pixel in the first source image can be calculated by combining the center point position of each image window with the local homography matrix corresponding to that image window, using methods such as bilinear interpolation.

[0084] Based on the above scheme, optionally, determining the target homography matrix corresponding to the first pixel based on the local homography matrix corresponding to the image window where the first pixel is located includes: determining the image window where the first pixel is located; if there are multiple image windows, determining the weight of the image window based on the distance of the first pixel from the center point of the image window; and performing weighted fusion on the local homography matrix corresponding to the image window based on the weight of the image window to obtain the target homography matrix corresponding to the first pixel.

[0085] Specifically, for any first pixel in the first source image, the multiple image windows in which it is located due to window overlap can be determined; then, according to the preset weighting rules, such as assigning weights inversely proportional to the distance from the first pixel to the center point of each window, the weight value corresponding to each target window is calculated; then, the local homography matrix of each target window is weighted and fused according to the weight value to obtain the final transformation matrix of the first pixel for pixel mapping, that is, the target homography matrix.

[0086] The target homography matrix can be calculated based on the following formula:

[0087] ;

[0088] in, The first pixel The corresponding target homography matrix; The first pixel The image window The corresponding local homography matrix; It is an image window The center point; Represents the first pixel To the center of the image window The distance; This represents a weighting function, such as a normalization function.

[0089] By assigning weights based on the distance from the pixel to the center of the window when the first pixel is in multiple windows, the window with the more accurate local feature characterization of the pixel has a higher weight. As a result, the fused target homography matrix can better match the actual local transformation features of the pixel. Based on the weighted complementarity of the local matrices of multiple windows, the mapping deviation caused by the lack of local features in a single window matrix can be compensated, further improving the reliability of the target homography matrix and providing more accurate support for pixel mapping.

[0090] Furthermore, the second pixel in the second source image can be determined based on the first pixel in the first source image and its corresponding target homography matrix, thereby achieving pixel-level accurate transformation mapping.

[0091] By determining the target homography matrix corresponding to the local homography matrix of the window where the first pixel is located, and then mapping the second pixel using this matrix, the mapping deviation caused by directly using the global matrix and ignoring local details is avoided. The calculated target homography matrix can achieve pixel-level accurate matching, effectively ensuring the stitching consistency of the two source images in local details, reducing problems such as edge misalignment and detail distortion, and improving the quality of the stitched image.

[0092] Optionally, based on the target homography transformation matrix, the first source image and the second source image can be resampled respectively. The pixels in the first source image and the second source image are mapped to the stitched image coordinate system with the upper left corner of the first source image as the origin according to the target homography transformation matrix corresponding to the pixels. According to the spatial correspondence between multiple first pixels in the first source image and their corresponding second pixels in the second source image, the pixels of the first source image and the second source image in the image group are fused and stitched to obtain the target image of the object region corresponding to the image group in the target object.

[0093] Before determining the second pose matrix of the camera in the second coordinate system based on the first pose matrix and the hand-eye conversion matrix, or before performing pixel coordinate correction on the first source image and the second source image in the image group based on the camera intrinsic matrix and camera distortion parameters of the camera, the method may optionally further include: taking pictures of a preset calibration board at different poses of the end of the robotic arm using a camera set at the end of the robotic arm to obtain multiple calibration board images; and determining the hand-eye conversion matrix, the camera intrinsic matrix, and the camera distortion parameters corresponding to the end of the robotic arm based on the multiple calibration board images.

[0094] The preset calibration board can be a standard board with a known pattern and size. The board surface can be printed with regular grids, dots, and other patterns, and the actual coordinates of each feature point in the pattern are known.

[0095] Optionally, when calibrating the hand-eye conversion matrix, camera intrinsic parameter matrix, and camera distortion parameters, multiple images of the preset calibration board can be captured by the camera. These calibration board images can be obtained by capturing images of the preset calibration board with the camera in different poses at the end effector of the robotic arm.

[0096] Optionally, when photographing the calibration plate, by using methods such as Zhang Zhengyou's calibration method, controlling the movement of the robotic arm, and adjusting the position of the robotic arm's end effector, images of the calibration plate at different angles and positions can be obtained, thereby improving the accuracy of the calibration parameters.

[0097] Furthermore, the hand-eye transformation matrix corresponding to the end effector of the robotic arm, the camera intrinsic parameter matrix, and the camera distortion parameters can be obtained by calculating the pixel coordinates and actual coordinates of feature points in multiple calibration plate images.

[0098] Multiple images of the calibration board are obtained by taking pictures of the calibration board in different poses by a robotic arm. Based on the known precise coordinate information of the calibration board, the hand-eye transformation matrix, camera intrinsic parameter matrix and distortion parameters can be accurately solved, avoiding the impact of unknown parameters or estimation errors on subsequent pose calculation and pixel correction. The acquisition of multi-view calibration data can improve the robustness of parameter calibration and improve the accuracy and reliability of image stitching.

[0099] The technical solution of this invention first involves acquiring multiple source images of a target object using a camera mounted at the end of a robotic arm, determining a global transformation matrix between the source images, and defining multiple image groups based on the source images. Each image group includes a first source image and a second source image to be stitched together, with some image regions in the first and second source images corresponding to the same object region of the target object. The precise movement of the robotic arm controls the camera's shooting angle and position, ensuring stable viewing angles and controllable overlapping areas in the acquired source images, laying a high-quality data foundation for subsequent stitching. Determining the global transformation matrix between the multiple source images establishes an overall spatial relationship between the images, reducing the accumulation of errors in subsequent local processing. Next, for each image group, multiple target feature point pairs in the first and second source images within the image group are determined. The first and second source images are then divided into multiple image windows with overlapping areas, and the local homography corresponding to each image window is determined based on the target feature point pairs in the image window and the global transformation matrix. The homography matrix can improve the matching accuracy of the correspondence between images by using multiple target feature points. It divides image windows with overlapping regions, calculates the local homography matrix corresponding to each image window, and accurately adapts to subtle local changes within the window while following global transformation rules, thus improving the accuracy of the transformation matrix in describing the spatial relationship of the image. Finally, for the first pixel in the first source image, the second pixel in the second source image is determined based on the local homography matrix associated with the first pixel. The first source image and the second source image are then stitched together based on multiple first pixels and their corresponding second pixels to obtain the target image of the target object. This achieves pixel-level accurate spatial mapping by stitching images together using multiple sets of spatial correspondences between first pixels and corresponding second pixels. This reduces stitching errors through mutual verification of a large amount of pixel-level data, ensuring the overall structural integrity and local detail clarity of the generated target image. By balancing accuracy and efficiency through a transformation matrix calculation that combines global and local approaches, and achieving high-quality stitching with pixel-level mapping, the target image of the target object can be accurately presented.

[0100] Example 2

[0101] Figure 2This is a flowchart of an image stitching method provided in Embodiment 2 of the present invention. Based on the above optional technical solutions, optionally, determining multiple target feature point pairs in the first source image and the second source image in the image group includes: determining multiple target feature point pairs in the first source image and the second source image based on the first source image, the second source image, and an image detection model; wherein, the image detection model is obtained by training a machine learning model based on sample image pairs and their corresponding multiple expected feature point pairs. For specific implementation details, please refer to the description of this embodiment. Technical features that are the same as or similar to those in the foregoing embodiments will not be repeated here. Figure 2 As shown, the method may specifically include:

[0102] S210. Multiple source images of the target object are acquired by a camera set at the end of the robotic arm, a global transformation matrix between the multiple source images is determined, and multiple image groups are determined based on the multiple source images. The image groups include a first source image and a second source image to be stitched together, and some image regions in the first source image and the second source image correspond to the same object region of the target object.

[0103] S220. For each image group, determine multiple target feature point pairs in the first source image and the second source image based on the first source image, the second source image and the image detection model.

[0104] Among them, the image detection model can be understood as a model trained by a machine learning model that can process the input first source image and second source image, extract feature points in the image and complete feature point matching, thereby determining the target feature point pair, which can provide key feature matching support for image stitching.

[0105] The image detection model is trained on a machine learning model based on sample image pairs and their corresponding multiple expected feature point pairs. A sample image pair refers to a pair of images used to train the image detection model, consisting of two images that have overlapping regions and correspond to the same object region. Expected feature point pairs can be manually labeled or otherwise determined matching feature point pairs from the sample image pairs.

[0106] Optionally, a large number of sample image pairs can be collected, and the expected feature point pairs that match each other in each sample image pair can be labeled. Then, the sample image pairs are input into the machine learning model, and the expected feature point pairs are used as labels to correct and optimize the feature point matching results output by the model. This allows the model to gradually learn the feature point matching rules between images, and finally trains an image detection model that can accurately extract target feature point pairs in image pairs.

[0107] The image detection model can be composed of a feature extraction module, a description generation module, and a feature matching module. The feature extraction module receives input first and second source images and extracts representative image features to provide basic feature information for subsequent feature descriptor generation. The description generation module further processes the first and second image features obtained by the feature extraction module to generate corresponding feature descriptors for multiple feature points in each source image, which can be used for subsequent feature point matching. The feature matching module is the functional module in the image detection model that performs feature point matching. The feature matching module can calculate and compare the similarity between the first and second feature descriptors to filter out mutually matching feature points, thereby obtaining multiple target feature point pairs in the first and second source images.

[0108] Based on the above scheme, optionally, determining multiple target feature point pairs in the first source image and the second source image according to the first source image, the second source image, and the image detection model includes: inputting the first source image and the second source image into the feature extraction module of the image detection model to obtain first image features of the first source image and second image features of the second source image; inputting the first image features and the second image features into the description generation module of the image detection model to obtain first feature descriptors of multiple first feature points of the first source image and second feature descriptors of multiple second feature points of the second source image; and inputting the first feature descriptors and the second feature descriptors into the feature matching module of the image detection model to obtain multiple target feature point pairs in the first source image and the second source image.

[0109] Specifically, the first source image and the second source image can be input into the feature extraction module. The convolutional neural network in the feature extraction module performs layer-by-layer convolution operations on the input first and second source images, thereby extracting image features at different scales and semantic levels, resulting in first image features of the first source image and second image features of the second source image. The first image feature refers to the information that reflects the attribute characteristics of the first source image after feature extraction by the feature extraction module. For example, image edges, local structures, and texture features can be used as the basis for generating the first feature descriptor. The second image feature refers to the information that reflects the attribute characteristics of the second source image after feature extraction by the feature extraction module. It corresponds to the first image feature and can be used as the basis for generating the second feature descriptor.

[0110] Specifically, the description generation module can generate highly discriminative feature descriptors based on the extracted first and second image features, combined with key feature information such as the position, scale, and orientation of the first and second feature points. These descriptors consist of first feature descriptors for multiple first feature points in the first source image and second feature descriptors for multiple second feature points in the second source image, representing the uniqueness of each feature point. The first feature points can be pixels extracted from the first source image by the feature extraction module, possessing distinct features such as corner points or edge points. Each first feature point has a corresponding first feature descriptor for matching with feature points in the second source image. The first feature descriptor can be understood as a vector or data result obtained after quantifying the feature information of the first feature point. It uniquely represents the features of the first feature point, facilitating similarity comparison between the feature matching module and the second feature descriptor to achieve feature point matching. The second feature points refer to pixels with distinct features extracted from the second source image by the feature extraction module. The second feature points correspond to the first feature points. Each second feature point has a corresponding second feature descriptor for matching with feature points in the first source image. The second feature descriptor can be understood as a vector or data result obtained after quantifying the feature information of the second feature point. It uniquely represents the features of the second feature point and is key data for similarity comparison and feature point matching with the first feature descriptor.

[0111] Specifically, the first feature descriptor and the second feature descriptor can be input into the feature matching module, and similar feature points can be matched through the feature descriptors to obtain multiple target feature point pairs in the first source image and the second source image.

[0112] For example, taking source image 1 and source image 2 as examples, assume that the feature points detected in source image 1 and source image 2 are respectively and The matched target feature point pairs are ,in, , , representing feature points in source image 1 Feature points in source image 2 Matching.

[0113] Compared with traditional feature matching algorithms, such as Scale-Invariant Feature Transform (SIFT) and Speeded Up Robust Features (SURF), the feature matching method based on image detection models has superior overall performance. In complex scenes such as severe background interference and variable target shapes, its feature matching accuracy is higher, and it can more accurately establish the feature correspondence between different source images. At the same time, the method is faster in computation, which can effectively shorten the time spent on feature matching. In addition, it is more adaptable to changes in image illumination and viewing angle, and can still stably achieve accurate matching of feature points even when there are large differences in illumination intensity or significant shifts in shooting angle.

[0114] By integrating the independent yet collaborative functions of each module in the image detection model, the feature extraction module can accurately capture the key visual features of two source images, providing a solid foundation for subsequent matching; the description generation module can construct exclusive feature descriptors for feature points, enhancing the distinguishability of different feature points and reducing the risk of confusion due to similar features; the feature matching module can perform targeted descriptor comparison, quickly locking in matching relationships, realizing a standardized process from extraction to description to matching, reducing the impact of anomalies in a single module on the overall results, and can also optimize specific modules according to actual needs, such as improving feature extraction speed and optimizing matching algorithms, significantly enhancing the model's adaptability and flexibility to different scenarios.

[0115] S230. Divide the first source image and the second source image into multiple image windows with overlapping regions, and determine the local homography matrix corresponding to the image window based on the target feature point pairs in the image window and the global transformation matrix.

[0116] S240. For a first pixel in the first source image, determine a second pixel in the second source image based on the local homography matrix associated with the first pixel, and stitch the first source image and the second source image together based on multiple first pixels and their corresponding second pixels to obtain a target image of the target object.

[0117] The technical solution of this invention determines multiple target feature point pairs in the first source image and the second source image based on the first source image, the second source image, and an image detection model. The image detection model is trained on a machine learning model using sample image pairs and their corresponding multiple expected feature point pairs. The image detection model can accurately identify feature points in the first and second source images, reducing mismatches caused by background interference and lighting changes, ensuring the reliability of target feature point pairs. After training, the model can automatically complete feature point detection and matching without manual intervention in feature point selection, significantly shortening the time required to obtain target feature point pairs from two source images. This significantly reduces operating costs, achieves efficient stitching, and allows the integration of sample image pairs from different scenarios during training, enabling the image detection model to have strong generalization ability. Even if the first and second source images have differences in shooting conditions such as uneven lighting or viewing angle shifts, the model can still stably extract target feature point pairs, improving the practicality of image stitching.

[0118] Example 3

[0119] As an optional example of an embodiment of the present invention, such as Figure 3a As shown, the image stitching method may specifically include:

[0120] S310, Global Image Registration Based on Robotic Arm Pose

[0121] Since the parallax in the source image group is large, global alignment and local alignment can be performed in stages. In S310, the global projection matrix is ​​obtained based on the existing information, and the images are projected onto the same plane.

[0122] S311. Camera Calibration and Hand-Eye Calibration. The hand-eye calibration method is used to determine the transformation relationship between the robot arm's end-effector coordinate system E and the camera coordinate system C. Methods such as Zhang Zhengyou's calibration method can be used. By controlling the movement of the robotic arm, the camera can capture images of the calibration board in different poses. The hand-eye transformation matrix can then be solved using image processing algorithms. Camera intrinsic parameter matrix K, camera distortion parameter S.

[0123] S312. Acquire the source image to be stitched.

[0124] Based on the target being photographed, a pre-set shooting position is established, the robotic arm is moved, and the source image is captured using a camera at the end of the robotic arm.

[0125] Let the pose of the robotic arm's end effector in the base coordinate system be . From the position of the robotic arm end effector Relationship with hand-eye transition The pose of the camera in the robot arm's base coordinate system is calculated through coordinate transformation. Through matrix multiplication The pose matrix of the camera in the reference coordinate system can be obtained. This matrix contains the camera's position information (translation part) and attitude information (rotation part). The 3×3 submatrix in the upper left corner is the rotation matrix R describing the camera's attitude, and the first three elements of the fourth column constitute the translation matrix T describing the camera's position.

[0126] ;

[0127] S313. Calculate the global projection matrix of the image. For the source image acquired in S312, denote the corresponding camera pose matrix as follows: Let the transformation matrix between two images be denoted. Assuming points in source image 1 The corresponding point in source image 2 is Then there is .in, .

[0128] S320, Local fine-grained alignment based on semantic feature points

[0129] After obtaining the global transformation matrix of S310, the local mismatch caused by disparity is further fine-tuned based on the improved APAP algorithm. First, semantic feature points are obtained using an image detection model, then the obtained feature point pairs are filtered, and finally, image alignment and fusion are achieved based on the Moving Direct Linear Transform (MDLT) and the global transformation matrix.

[0130] S321. Image Preprocessing. Based on the camera intrinsic parameter matrix K and camera distortion parameters S obtained in S310, the first and second source images are processed to aid in the accurate extraction and matching of subsequent features. According to the principle of image correction, a mathematical model is established using the camera intrinsic parameter matrix K and distortion parameters S to recalculate the coordinates of each pixel in the first and second source images, eliminating the distortion caused by lens distortion by following the inverse process of image distortion. Subsequently, histogram analysis of brightness distribution is used, combined with grayscale correction methods to correct the brightness difference between the first and second source images.

[0131] S322, Feature Detection and Matching. The image detection model mainly consists of a feature extraction module, a description generation module, and a feature matching module. First, the first and second source images to be stitched together are... and The input is fed into the feature extraction module, where the convolutional neural network processes the input source image. and Layer-by-layer convolution operations are performed to extract image features at different scales and semantic levels. These features effectively represent the local structure and texture information of the image. Subsequently, in the description generation module, the model generates highly discriminative feature descriptors based on the extracted features and key information such as the location, scale, and orientation of feature points, thus representing the uniqueness of each feature point. Finally, in the feature matching module, the model quickly and accurately finds the correspondence between feature points in two images by calculating the similarity between different image feature descriptors. During the calculation process, the model generates descriptors based on the location, scale, and orientation of feature points to represent their uniqueness, and uses these descriptors to find similar feature points, thereby determining matching feature point pairs. (Source image...) and The detected feature points are respectively and The matched feature point pairs are Compared with traditional algorithms such as Scale Invariant Feature Transform (SIFT) and Speeded Robust Feature Transform (SURF), model-based feature matching has higher matching accuracy in complex scenes, faster computation speed, and can adapt to images with different lighting and viewpoints.

[0132] S323, Adaptive window partitioning and local transformation model estimation. The transformation matrix obtained from S310 is used. As the initial local homography matrix The image is divided into multiple overlapping small windows, each with a center point. The size and overlap of the windows are adjustable parameters. Smaller windows can better capture local deformations but increase computational cost; larger windows are more computationally efficient but may not accurately reflect complex local changes. Assume the window size is... The overlap rate is The image is divided into windows with a certain step size (determined by the window size and overlap ratio). For each small window, a local homography matrix is ​​estimated using the least squares method based on the feature points within the window. Let the feature point pairs within the window be... , Given the number of feature point pairs within the window, the objective function is minimized by solving the problem. get ,in It represents Euclidean distance.

[0133] S324, transform resampling and stitching fusion.

[0134] Based on the center point position of each window and the estimated final local homography matrix, the final transformation is calculated for each pixel in the entire image using methods such as bilinear interpolation. For any point in the image... Find the multiple windows containing it. (Due to window overlap, a point may lie in multiple windows.) Based on the window weights, the homography matrices of these windows are weighted and fused to obtain the point. The final transformation matrix to determine the point Position in the stitched image.

[0135] ;

[0136] in For window The corresponding transformation matrix, It is a window center point Point arrive distance, This represents a weighting function, such as a normalization function.

[0137] Using the transformation relationship obtained above, the image is resampled, and the pixels in the image are mapped to the coordinate system of the stitched image (generally with the upper left corner of the first source image as the origin) to obtain the final stitched image.

[0138] The technical solution of this invention first calculates the camera pose matrix based on the pose in the coordinate system of the robotic arm's end effector and the hand-eye transformation matrix. Then, it calculates the global projection matrix of the image based on the camera pose matrix, ensuring the accuracy of the camera pose matrix and providing a reliable foundation for the overall spatial alignment of subsequent image stitching. Next, it preprocesses the image using an image detection model, extracting matching feature point pairs from the source image. The image is then divided into windows with overlapping regions, and the local homography matrix of each window is calculated, improving the accuracy of subsequent feature point extraction. Extracting matching feature points through the model effectively reduces the false matching rate and provides high-quality data for calculating the local homography matrix. The system supports the calculation of local homography matrices based on feature point pairs within a window, which can accurately characterize subtle changes in each region, compensate for the shortcomings of global matrices in ignoring local differences, and improve the accuracy of transformation description. Finally, the final transformation matrix of each pixel is determined based on the local homography matrix of each window, and then the first and second source images are stitched together to generate the target image, achieving pixel-level accurate spatial mapping. This avoids the local detail distortion caused by traditional global transformations. By completing the stitching of two source images based on pixel-level transformation matrices, the stitching error can be further reduced through the constraint of a large amount of pixel-level data, ensuring that the target image achieves uniformity in overall structural integrity and local detail clarity, and generating a high-quality target image.

[0139] Example 4

[0140] Figure 4This is a schematic diagram of an image stitching device according to Embodiment 4 of the present invention. This device is used to execute the image stitching method provided in any of the above embodiments. This device and the image stitching methods of the above embodiments belong to the same inventive concept. Details not described in detail in the embodiments of the image stitching device can be found in the embodiments of the above image stitching methods. Figure 4 As shown, the device includes: an image group determination module 410, a window local homography matrix determination module 420, and a target image acquisition module 430.

[0141] The image group determination module 410 is used to acquire multiple source images of the target object using a camera set at the end of the robotic arm, determine a global transformation matrix between the multiple source images, and determine multiple image groups based on the multiple source images. Each image group includes a first source image and a second source image to be stitched together, and some image regions in the first and second source images correspond to the same object region of the target object. The window local homography matrix determination module 420 is used to determine multiple target feature point pairs in the first and second source images of each image group, and to divide the first and second source images into multiple image windows with overlapping regions, and to determine the local homography matrix corresponding to the image window based on the target feature point pairs in the image window and the global transformation matrix. The target image acquisition module 430 is used to determine a second pixel in the second source image based on the local homography matrix associated with the first pixel in the first source image, and to stitch the first and second source images together based on the multiple first pixels and their corresponding second pixels to obtain a target image of the target object.

[0142] The technical solution of this invention firstly involves an image group determination module 410, which uses a camera mounted at the end of a robotic arm to acquire multiple source images of a target object. A global transformation matrix is ​​then determined between these source images. Multiple image groups are determined based on the source images, where each image group includes a first source image and a second source image to be stitched together. Partial image regions in the first and second source images correspond to the same object region of the target object. The precise movement of the robotic arm controls the camera's shooting angle and position, ensuring stable viewing angles and controllable overlapping areas in the acquired source images. This lays a high-quality data foundation for subsequent stitching. The global transformation matrix between the multiple source images is determined, establishing an overall spatial relationship between the images and reducing the accumulation of errors in subsequent local processing. Next, a window local homography matrix determination module 420, for each image group, determines multiple target feature point pairs in the first and second source images within the image group. The first and second source images are then divided into multiple image windows with overlapping areas. The image window is determined based on the target feature point pairs in the image window and the global transformation matrix. The local homography matrix corresponding to the target image is used to improve the matching accuracy of the correspondence between images by using multiple target feature points. Image windows with overlapping regions are divided, and the local homography matrix corresponding to each image window is calculated. While following the global transformation rules, it accurately adapts to subtle local transformations within the window, improving the accuracy of the transformation matrix in describing the spatial relationship of the image. Finally, the target image acquisition module 430 determines the second pixel in the second source image based on the local homography matrix associated with the first pixel in the first source image. The first source image and the second source image are then stitched together based on multiple first pixels and their corresponding second pixels to obtain the target image of the target object. This achieves pixel-level precise spatial mapping by using multiple sets of spatial correspondences between first pixels and corresponding second pixels for image stitching. This reduces stitching errors through mutual verification of a large amount of pixel-level data, ensuring the overall structural integrity and local detail clarity of the generated target image. By combining global and local transformation matrix calculations to balance accuracy and efficiency, and achieving high-quality stitching with pixel-level mapping, the target image of the target object can be accurately presented.

[0143] Based on the above scheme, optionally, the window local homography matrix determination module 420 includes a target feature point pair determination submodule. The target feature point pair determination submodule is used to determine multiple target feature point pairs in the first source image and the second source image based on the first source image, the second source image, and the image detection model; wherein the image detection model is obtained by training a machine learning model based on sample image pairs and their corresponding multiple expected feature point pairs.

[0144] Based on the above scheme, optionally, the image detection model includes a feature extraction module, a description generation module, and a feature matching module; the target feature point pair determination submodule includes an image feature acquisition unit, a feature descriptor acquisition unit, and a target feature point pair determination unit. Specifically, the image feature acquisition unit is used to input the first source image and the second source image into the feature extraction module of the image detection model to obtain first image features of the first source image and second image features of the second source image; the feature descriptor acquisition unit is used to input the first image features and the second image features into the description generation module of the image detection model to obtain first feature descriptors of multiple first feature points of the first source image and second feature descriptors of multiple second feature points of the second source image; the target feature point pair determination unit is used to input the first feature descriptors and the second feature descriptors into the feature matching module of the image detection model to obtain multiple target feature point pairs in the first source image and the second source image.

[0145] Based on the above scheme, optionally, the window local homography matrix determination module 420 includes a window local homography matrix determination submodule. This submodule is used to determine the local homography matrix corresponding to the image window based on the target feature point pairs in the image window and the least squares method, using the global transformation matrix as the initial homography matrix.

[0146] Based on the above scheme, optionally, the target image acquisition module 430 includes a second pixel point determination submodule. The second pixel point determination submodule is used to determine the target homography matrix corresponding to the first pixel point based on the local homography matrix corresponding to the image window where the first pixel point is located, and to determine the second pixel point in the second source image based on the first pixel point and the target homography matrix.

[0147] Based on the above scheme, optionally, the second pixel determination submodule includes an image window weight determination unit and a target homography matrix determination unit. The image window weight determination unit is used to determine the image window in which the first pixel is located. If there are multiple image windows, the weight of the image window is determined based on the distance of the first pixel from the center point of the image window. The target homography matrix determination unit is used to perform weighted fusion of the local homography matrix corresponding to the image window based on the weight of the image window to obtain the target homography matrix corresponding to the first pixel.

[0148] Based on the above scheme, optionally, the image group determination module 410 includes a pose matrix determination submodule and a global transformation matrix determination submodule. The pose matrix determination submodule is used to determine the first pose matrix of the robotic arm end effector in the first coordinate system, and to determine the second pose matrix of the camera in the second coordinate system based on the first pose matrix and the hand-eye transformation matrix; wherein the hand-eye transformation matrix is ​​used to indicate the transformation relationship between the first coordinate system and the second coordinate system; the global transformation matrix determination submodule is used to determine the global transformation matrix between the multiple source images based on the second pose matrices of the multiple source images.

[0149] Based on the above scheme, optionally, the window local homography matrix determination module 420 includes a pixel coordinate correction submodule, a brightness correction submodule, and a feature point determination submodule. The pixel coordinate correction submodule is used to perform pixel coordinate correction on the first source image and the second source image in the image group according to the camera intrinsic parameter matrix and camera distortion parameters of the camera, respectively. The brightness correction submodule is used to perform brightness correction on the first source image and / or the second source image after pixel coordinate correction, so as to reduce the brightness difference between the first source image and the second source image. The feature point determination submodule is used to determine multiple feature points in the first source image and the second source image after brightness correction.

[0150] Optionally, based on the above scheme, the device further includes a calibration module. The calibration module is used to capture images of a preset calibration board at different poses of the robotic arm end effector using a camera mounted on the end effector of the robotic arm, thereby obtaining multiple images of the calibration board; and to determine the hand-eye conversion matrix corresponding to the end effector of the robotic arm, the camera intrinsic parameter matrix of the camera, and camera distortion parameters based on the multiple calibration board images.

[0151] The image stitching device provided in the embodiments of the present invention can execute the image stitching method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0152] Example 5

[0153] Figure 5A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0154] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0155] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0156] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as image stitching methods.

[0157] In some embodiments, the image stitching method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the image stitching method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the image stitching method by any other suitable means (e.g., by means of firmware).

[0158] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0159] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0160] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0161] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0162] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0163] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0164] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from ROM 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of the embodiments of the present invention.

[0165] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0166] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. An image stitching method, characterized by, The method comprises the following steps: Collecting multiple source images of a target object by a camera arranged at the end of a mechanical arm, determining a global transformation matrix between the multiple source images, and determining multiple image groups according to the multiple source images, wherein the image groups include a first source image and a second source image to be spliced, and part of the image regions in the first source image and the second source image correspond to the same object region of the target object; For each image group, multiple target feature point pairs in the first source image and the second source image in the image group are determined, and the first source image and the second source image are respectively divided into multiple image windows with overlapping regions, and a local homography matrix corresponding to the image window is determined according to the target feature point pairs in the image window and the global transformation matrix; For a first pixel point in the first source image, a second pixel point in the second source image is determined according to the local homography matrix associated with the first pixel point, and the first source image and the second source image are spliced according to multiple first pixel points and their corresponding second pixel points to obtain a target image of the target object.

2. The image stitching method of claim 1, wherein, The determination of the multiple target feature point pairs in the first source image and the second source image in the image group comprises: Determining the multiple target feature point pairs in the first source image and the second source image according to the first source image, the second source image and an image detection model, wherein the image detection model is obtained by training a machine learning model according to a sample image pair and its corresponding multiple expected feature point pairs.

3. The image stitching method of claim 2, wherein, The image detection model comprises a feature extraction module, a description generation module and a feature matching module; the determination of the multiple target feature point pairs in the first source image and the second source image according to the first source image, the second source image and the image detection model comprises: Inputting the first source image and the second source image into the feature extraction module of the image detection model to obtain a first image feature of the first source image and a second image feature of the second source image; Inputting the first image feature and the second image feature into the description generation module of the image detection model to obtain a first feature descriptor of multiple first feature points of the first source image and a second feature descriptor of multiple second feature points of the second source image; Inputting the first feature descriptor and the second feature descriptor into the feature matching module of the image detection model to obtain the multiple target feature point pairs in the first source image and the second source image.

4. The image stitching method of claim 1, wherein, The determination of the local homography matrix corresponding to the image window according to the target feature point pairs in the image window and the global transformation matrix comprises: Taking the global transformation matrix as an initial homography matrix, and determining the local homography matrix corresponding to the image window according to the target feature point pairs in the image window and the least square method.

5. The image stitching method of claim 1, wherein, The determination of the second pixel point in the second source image according to the local homography matrix associated with the first pixel point comprises: Determine a target homography matrix corresponding to the first pixel point according to a local homography matrix corresponding to the image window in which the first pixel point is located.

6. The image stitching method of claim 5, wherein, The method further includes: Determine the image window in which the first pixel point is located, and determine a weight of the image window according to a distance between the first pixel point and a center point of the image window when there are multiple image windows; Weight and fuse the local homography matrices corresponding to the image windows according to the weights of the image windows to obtain the target homography matrix corresponding to the first pixel point.

7. The image stitching method of claim 1, wherein, The method further includes: Determine a first pose matrix of the end of the mechanical arm in a first coordinate system, and determine a second pose matrix of the camera in a second coordinate system according to the first pose matrix and a hand-eye conversion matrix, wherein the hand-eye conversion matrix is used to indicate a conversion relationship between the first coordinate system and the second coordinate system; Determine the global conversion matrix between the multiple source images according to the second pose matrices of the multiple source images.

8. The image stitching method of claim 1, wherein, The method further includes: Perform pixel coordinate correction on the first source image and the second source image in the image group according to a camera intrinsic matrix and camera distortion parameters of the camera; Perform brightness correction on the first source image and / or the second source image after the pixel coordinate correction to reduce a brightness difference between the first source image and the second source image; Determine multiple feature points in the first source image and the second source image after the brightness correction.

9. The image stitching method of claim 7 or 8, characterized in that, The method further includes: Capture a preset calibration board when the end of the mechanical arm is in different poses to obtain multiple calibration board images through a camera arranged at the end of the mechanical arm; Determine a hand-eye conversion matrix corresponding to the end of the mechanical arm, a camera intrinsic matrix, and camera distortion parameters of the camera according to the multiple calibration board images.

10. An image stitching apparatus characterized by comprising: The method further includes: An image group determination module is configured to collect multiple source images of a target object through a camera arranged at the end of a mechanical arm, determine a global conversion matrix between the multiple source images, and determine multiple image groups according to the multiple source images, wherein the image groups include a first source image and a second source image to be stitched, and part of image regions in the first source image and the second source image correspond to a same object region of the target object; A window local homography matrix determination module is configured to determine multiple target feature point pairs in the first source image and the second source image in each image group, divide the first source image and the second source image into multiple image windows with overlapping regions respectively, and determine a local homography matrix corresponding to each image window according to the target feature point pairs in the image window and the global conversion matrix. The target image acquisition module is configured to, for a first pixel point in the first source image, determine a second pixel point in the second source image according to the local homography matrix associated with the first pixel point, and stitch the first source image and the second source image according to a plurality of first pixel points and corresponding second pixel points thereof to obtain a target image of the target object.

11. An electronic device, comprising: Comprise: at least one processor; and a memory connected with the at least one processor in communication; wherein, the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the image stitching method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the image stitching method according to any one of claims 1-9 when executed.

13. A computer program product comprising a computer program, characterized in that, The computer program, when executed by the processor, implements the image stitching method according to any one of claims 1-9.