Image stitching method and device, electronic equipment and storage medium

CN122139210APending Publication Date: 2026-06-02SHENZHEN DANALE TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN DANALE TECH
Filing Date
2023-12-28
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Due to parallax problems when stitching images, the multi-eye camera device cannot guarantee that the object to be measured is clear and complete in the image after stitching, resulting in poor image stitching quality.

Method used

By acquiring images taken by different cameras in a multi-photo camera device at the same time, determining the target object in the overlapping area, and determining the target feature point filtering threshold based on the target object's characteristic information, focusing on the feature points of the target object with a high degree of importance for image stitching.

Benefits of technology

The quality of image stitching, especially the stitching quality of important areas, reduce stitching deviation, and achieve a clearer image stitching effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122139210A_ABST
    Figure CN122139210A_ABST
Patent Text Reader

Abstract

This application discloses an image stitching method, apparatus, electronic device, and storage medium. The method may include: acquiring a first image and a second image captured simultaneously by a first camera and a second camera, respectively, wherein the first camera and the second camera are different cameras in a multi-camera device, and the first image and the second image have an overlapping area; determining a target object in the overlapping area; determining a target feature point filtering threshold for the target object based on the target object's feature information, wherein the feature information indicates the importance of the target object, and the importance of the target object is negatively correlated with the target feature point filtering threshold; and stitching the first image and the second image together based on the feature points determined by the target feature point filtering threshold. Implementing this application embodiment can improve the accuracy of image stitching.
Need to check novelty before this filing date? Find Prior Art

Description

Image stitching method, device, electronic device and storage medium Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image stitching method, device, electronic device and storage medium. Background Art

[0002] A wide field of view is a crucial factor in evaluating camera performance. Fisheye cameras offer a wide field of view, but suffer from significant distortion. While pan-tilt cameras can rotate and cover a wider field of view, they cannot capture all environmental information at the same time. Simply combining images from cameras in different locations can lead to a lack of clear spatial perception.

[0003] Multi-camera systems overcome the limitations of single cameras to a certain extent. Comprising two or more cameras, they collaborate to achieve a wider capture range and more viewing angles. However, due to parallax issues when stitching images, multi-camera systems cannot guarantee that the object being measured appears clear and complete in the stitched image, resulting in poor image quality.

[0004] Summary of the Invention

[0005] The embodiments of the present application disclose an image stitching method, device, electronic device, and storage medium, which can improve the quality of image stitching.

[0006] An embodiment of the present application discloses an image stitching method, the method comprising: acquiring a first image and a second image respectively captured by a first camera and a second camera at the same time, the first camera and the second camera being different cameras in a multi-camera device, and the first image and the second image having an overlapping area; determining a target object in the overlapping area; determining a target feature point screening threshold of the target object based on feature information of the target object, the feature information indicating the importance of the target object, and the importance of the target object being negatively correlated with the target feature point screening threshold; and stitching the first image and the second image based on feature points determined by the target feature point screening threshold.

[0007] As an optional implementation manner, the characteristic information of the target object includes one or more of the following: depth information, size of the recognition frame, confidence of the recognition frame, and object type.

[0008] The above-mentioned embodiment can comprehensively and accurately determine the importance of the target object based on one or more feature information of the depth information, the size of the recognition frame, the confidence of the recognition frame and the object type, so that the image stitching can be focused on the feature points of the target objects with high importance, which helps to perform image stitching based on important areas in the image, so that the stitching quality of the target objects with high importance in the stitched image is high, thereby improving the quality of image stitching.

[0009] As an optional implementation manner, the characteristic information of the target object includes depth information, and the overlapping area includes the first target object and the second target object; and determining the target feature point screening threshold of the target object based on the characteristic information of the target object includes:

[0010] Based on the first depth information of the first target object, a first target feature point screening threshold of the first target object is determined, and based on the second depth information of the second target object, a second target feature point screening threshold of the second target object is determined; the first depth information is a first distance between the first target object and a baseline, and the second depth information is a second distance between the second target object and the baseline, and the baseline is a line connecting the optical center of the first camera and the optical center of the second camera; if the first distance is smaller than the second distance, the first target feature point screening threshold is smaller than the second target feature point screening threshold.

[0011] In the above embodiment, the depth information of different target objects is different, and the importance of the target object can be determined based on the depth information. This helps to perform image stitching on the first image and the second image, mainly based on the feature points of the target object that is closer to the multi-camera device. This can improve the quality of image stitching and greatly improve the stitching quality of the area where the target object with small depth information is located.

[0012] As an optional implementation manner, the characteristic information of the target object includes at least one of the following: depth information, identification frame size, identification frame confidence, and object type. Before determining the feature point screening threshold of the target object based on the characteristic information of the target object, the method further includes:

[0013] Determining target detection information of the target object; the target detection information includes a first identification box corresponding to the target object in the first image, a second identification box corresponding to the target object in the second image, a first confidence level corresponding to the first identification box, and a second confidence level corresponding to the second identification box; the first confidence level is used to indicate the degree of confidence that the first identification box contains the target object, and the second confidence level is used to indicate the degree of confidence that the second identification box contains the target object;

[0014] Feature information of the target object is determined according to the target detection information of the target object and the depth information of the target object.

[0015] The above-mentioned embodiment combines target detection information and depth information to comprehensively and accurately determine the importance of the target object, thereby focusing on image stitching based on the feature points of the target object with high importance, which helps to perform image stitching based on important areas in the image, thereby improving the quality of image stitching.

[0016] As an optional implementation, the method of determining the target feature point screening threshold of the target object based on the feature information of the target object includes: determining the average area of ​​the identification frame corresponding to the target object based on the size of the first identification frame and the size of the second identification frame; determining the average confidence value corresponding to the target object based on the first confidence level and the second confidence level; and determining the target feature point screening threshold of the target object based on the average area of ​​the identification frame, the average confidence value and the depth information of the target object.

[0017] In the above embodiment, the average recognition box area, average confidence level and depth information of different target objects are different, and thus the target feature point screening thresholds of different target objects are different. The feature points extracted for different target objects are different. Therefore, the stitching quality of important target objects can be strategically optimized during image stitching.

[0018] As an optional implementation, the target feature point screening threshold of the target object is negatively correlated with the average area of ​​the identification frame, the target feature point screening threshold of the target object is negatively correlated with the average confidence value, and the target feature point screening threshold of the target object is positively correlated with the depth information.

[0019] In the above embodiment, the target object has a high average confidence value, a large average identification box area, and small depth information. The corresponding target object is close in the image, has a large area, and is highly important. Therefore, setting a smaller target feature point screening threshold can retain more feature points, which helps to consider the target object more when stitching images, so as to achieve higher quality stitching.

[0020] As an optional implementation, the target feature point screening threshold of the target object is determined based on the average area value of the identification frame, the average confidence value and the depth information of the target object, including: performing weighted summation calculation on the average area value of the identification frame, the average confidence value and the depth information of the target object according to the weight values ​​corresponding to the average area value of the identification frame, the average confidence value and the depth information of the target object, and determining the target feature point screening threshold of the target object according to the weighted summation calculation result; wherein, among the weight values ​​corresponding to the average area value of the identification frame, the average confidence value and the depth information of the target object, the absolute value of the weight value corresponding to the depth information is the largest.

[0021] In the above embodiment, when the target feature point screening threshold of the target object is determined by weighted summation calculation, among the weight values ​​corresponding to the average value of the identification box area, the average value of the confidence and the depth information of the target object, the absolute value of the weight value corresponding to the depth information is the largest. This is because the depth information has a greater impact on the importance of the target object. If the target object is dynamic and moves in a direction that is not parallel to the baseline, the depth information of the target object will change continuously. The depth information is one of the information that can most sensitively and timely reflect the dynamic changes of the target object. Therefore, the weight value corresponding to the depth information of the target object is set to the maximum, which can pay more attention to the splicing effect of the moving target object in the overlapping area.

[0022] As an optional implementation, the method of determining the target feature point screening threshold of the target object based on the weighted summation calculation result includes: normalizing the weighted summation calculation result; determining an adjustment coefficient based on the weighted summation calculation result after normalization, wherein the adjustment coefficient is a positive number less than or equal to 1; and determining the target feature point screening threshold of the target object based on the preset feature point screening threshold and the adjustment coefficient.

[0023] In the above embodiment, the adjustment coefficient is determined based on the weighted sum calculation result after normalization processing, and the preset feature point screening threshold is adjusted, so that the adjustment coefficient is determined in real time based on the average value of the recognition frame area, the average value of the confidence level and the depth information of the target object, so as to dynamically adjust the preset feature point screening threshold, so as to flexibly and accurately determine the target feature point screening threshold of the target object.

[0024] As an optional embodiment, before stitching the first image and the second image based on the feature points determined by the target feature point screening threshold, the method further includes: extracting feature points located within the first identification frame in the first image based on the target feature point screening threshold, and extracting feature points located outside the first identification frame in the first image based on a preset feature point screening threshold; determining the feature points within the first identification frame and the feature points outside the first identification frame as first feature points corresponding to the target object in the first image; extracting feature points located within the second identification frame in the second image based on the target feature point screening threshold, and extracting feature points located outside the second identification frame in the second image based on the preset feature point screening threshold; determining the feature points within the second identification frame and the feature points outside the second identification frame as second feature points corresponding to the target object in the second image.

[0025] In the above embodiment, feature points within each identification box can be detected based on the target feature point screening threshold, and feature points outside the identification box in the image can be detected based on the preset feature point screening threshold. The feature points extracted from the area where the target object is located are more than the feature points extracted from the background area, thereby focusing on solving the stitching quality of the important target object area, and taking into account the stitching quality of the background area, thereby comprehensively improving the quality of image stitching.

[0026] As an optional embodiment, the image stitching of the first image and the second image based on the feature points determined by the target feature point screening threshold includes: matching the first feature points corresponding to the target object in the first image with the second feature points corresponding to the target object in the second image to stitch the first image and the second image.

[0027] In the above embodiment, by matching the feature points determined by the first image and the second image, the mapping relationship between the two images can be calculated, and image stitching can be performed based on the mapping relationship between the two images, thereby reducing the stitching deviation caused by differences in camera position, angle or scale.

[0028] As an optional implementation, the matching of the first feature point corresponding to the target object in the first image and the second feature point corresponding to the target object in the second image to perform image stitching on the first image and the second image includes: matching the first feature point corresponding to the target object in the first image with the second feature point corresponding to the target object in the second image to obtain a homography matrix of the second image relative to the first image; performing a perspective transformation on the second image according to the homography matrix to map the second image to the camera coordinate system corresponding to the first image; and aligning the pixel points corresponding to the same coordinate values ​​in the first image and the transformed second image to perform image stitching on the first image and the second image.

[0029] In the above embodiment, by calculating the homography matrix of the entire second image relative to the first image, only the same homography matrix is ​​used for transformation, the seam distribution is simple, and seam optimization is convenient, so that the seams of image stitching can transition smoothly, thereby solving the problem of obvious seams and improving the image stitching effect.

[0030] As an optional embodiment, after the first image and the second image are stitched together, the method further includes: determining the first pixel value of the pixel points of the target column in the first image and the second pixel value of the pixel points of the target column in the second image in the aligned area of ​​the first image and the second image, the aligned area being the area where the pixels in the first image and the second image are aligned; the pixel points of the target column are any column of pixel points in the aligned area; performing weighted fusion on the pixel points of the target column according to the first weight and the first pixel value of the pixel points of the target column, and according to the second weight and the second pixel value of the pixel points of the target column to obtain the target pixel value corresponding to the pixel points of the target column.

[0031] In the above embodiment, weighted fusion processing is performed on the aligned regions in the image stitching result, so that the transition of the image stitching is smoother.

[0032] As an optional embodiment, before weighted fusion of the pixel points of the target column according to the first weight and the first pixel value of the pixel points of the target column, and according to the second weight and the second pixel value of the pixel points of the target column, the method further includes: determining a channel score according to the difference in the number of columns between the target column and the leftmost column in the alignment area, and the difference in the number of columns between the leftmost column and the rightmost column in the alignment area; and determining the first weight and the second weight according to the channel score.

[0033] In the above embodiment, weighted fusion processing is performed on the aligned regions in the image stitching result, and the first weight and the second weight are determined according to the channel scores, thereby improving the accuracy of the weighted fusion and making the transition of the image stitching smoother.

[0034] As an optional embodiment, the characteristic information of the target object includes an object category; determining the target feature point screening threshold of the target object based on the characteristic information of the target object includes: judging whether the object category of the target object meets the preset object category; if the object category of the target object meets the preset object category, determining the third target feature point screening threshold of the target object; if the object category of the target object does not meet the preset object category, determining the fourth target feature point screening threshold of the target object; the third target feature point screening threshold is less than the fourth target feature point screening threshold.

[0035] In the above embodiment, the preset object category may be an object category of interest preset by the user. Determining the target feature point screening threshold according to the object category helps to improve the stitching quality of the area where the target object of interest is located.

[0036] As an optional embodiment, after obtaining the first image and the second image respectively taken by the first camera and the second camera at the same time, the method also includes: performing distortion correction on the first image and the second image according to the internal parameters of the first camera and the second camera; and / or, performing perspective transformation on the first image and the second image according to the internal parameters and external parameters of the first camera and the second camera; and / or, performing brightness balance on the first image and the second image.

[0037] In the above embodiment, the brightness of the first image and the second image are unified, and the parallax caused by camera rotation is solved through distortion correction and perspective transformation, so as to ensure that the same target object is imaged consistently in different images, which helps to improve the quality of image stitching.

[0038] An embodiment of the present application discloses an image stitching device, which includes: an image acquisition module for acquiring a first image and a second image respectively taken by a first camera and a second camera at the same time, wherein the first camera and the second camera are different cameras in a multi-camera device, and there is an overlapping area between the first image and the second image; an object determination module for determining a target object in the overlapping area; a threshold determination module for determining a target feature point screening threshold of the target object based on feature information of the target object, wherein the feature information indicates the importance of the target object, and the importance of the target object is negatively correlated with the target feature point screening threshold; and an image stitching module for stitching the first image and the second image based on feature points determined by the target feature point screening threshold.

[0039] An embodiment of the present application discloses an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor implements any one of the image stitching methods disclosed in the embodiments of the present application.

[0040] The present application discloses a multi-camera device, including a memory, a processor, and a multi-camera:

[0041] The multi-camera is used to capture multiple images containing overlapping areas. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the one or more processors, the one or more processors execute any one of the image stitching methods disclosed in the embodiments of the present application to achieve stitching of the multiple images.

[0042] The embodiments of the present application disclose one or more non-volatile computer-readable storage media storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute any one of the image stitching methods disclosed in the embodiments of the present application.

[0043] Compared with the related art, the embodiments of the present application have the following beneficial effects:

[0044] Acquire a first image and a second image captured at the same time by a first camera and a second camera, respectively, wherein the first camera and the second camera are different cameras in a multi-camera device, and there is an overlapping area between the first image and the second image; determine a target object in the overlapping area, and determine a target feature point screening threshold for the target object based on feature information of the target object, wherein the feature information can indicate the importance of the target object, and the importance of the target object is negatively correlated with the target feature point screening threshold; and perform image stitching on the first image and the second image based on the feature points determined by the target feature point screening threshold. In an embodiment of the present application, determine a target object in the overlapping area between the first image captured by the first camera and the second image captured by the second camera, and determine a target feature point screening threshold for the target object based on the feature information of the target object, and perform image stitching based on the feature points determined by the target feature point screening threshold. Since the feature information of the target object can reflect the importance of the target object, determining the target feature point screening threshold for the target object based on the importance of the target object can focus on performing image stitching based on feature points of target objects with high importance, which helps to perform image stitching based on important areas in the image, thereby improving the quality of image stitching. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0046] FIG1 is a schematic diagram of imaging of a multi-camera device in one embodiment;

[0047] FIG2 a is a diagram showing an application scenario of an image stitching method according to an embodiment;

[0048] FIG2 b is a schematic structural diagram of a multi-camera device according to an embodiment;

[0049] FIG3 is a schematic diagram of a flow chart of an image stitching method according to an embodiment;

[0050] FIG4 is a schematic diagram of perspective transformation in one embodiment;

[0051] FIG5 is a schematic diagram of a parallax method according to an embodiment;

[0052] FIG6 is a schematic flow chart of an image stitching method according to another embodiment;

[0053] FIG7 a is a schematic diagram of a target object in a first image and a second image in one embodiment;

[0054] FIG7 b is an image stitching result after the first image and the second image in FIG7 a are stitched together in one embodiment;

[0055] FIG7 c is a schematic diagram of a target object in a first image and a second image according to another embodiment;

[0056] FIG7 d is an image stitching result after the first image and the second image in FIG7 c are stitched together in one embodiment;

[0057] FIG8 is a schematic flow chart of an image stitching method according to another embodiment;

[0058] FIG9 is a schematic flow chart of an image stitching method according to another embodiment;

[0059] FIG10 is a schematic structural diagram of an image stitching device according to an embodiment;

[0060] FIG11 is a schematic structural diagram of an electronic device according to an embodiment;

[0061] FIG12 is a block diagram of a multi-camera device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0062] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0063] It should be noted that the terms "including," "having," and any variations thereof in the embodiments and drawings of this application are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.

[0064] It will be understood that the terms "first," "second," etc., used herein may be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, a first camera may be referred to as a second camera, and similarly, a second camera may be referred to as a first camera without departing from the scope of this application. The first camera and the second camera are both cameras, but they are not the same camera.

[0065] Binocular stereo vision is an important form of machine vision that estimates the distance of objects in a scene based on the parallax between two cameras. Binocular ranging simulates the principles of human vision, using computers to passively perceive distance. An object is viewed from two or more points, capturing images from different perspectives. Based on the matching pixel relationships between the images, the offset between the pixels is calculated using triangulation principles to obtain three-dimensional information about the object.

[0066] When stitching images together using multiple cameras, parallax can prevent objects in overlapping areas from appearing clear and complete in the stitched image. Parallax refers to the difference in direction when observing the same object from two points at a certain distance. The angle between two points as seen from the object is called the parallax angle, and the line connecting the two points is called the baseline. Knowing the parallax angle and the length of the baseline allows calculation of the distance between the object and the observer. Parallax is a visual term, also known as parallelism.

[0067] As shown in Figure 1, a schematic diagram of imaging in a multi-camera device according to one embodiment is shown. Let O1 and O2 be the optical centers of the two cameras in the multi-camera device, P1 and P2 be the locations of the two targets, and the focal length of the two cameras be f. The image points of P1 and P2 on the photoreceptor of the camera with O1 as the optical center are P1'1 and P2'1, respectively. The image points of P1 and P2 on the photoreceptor of the camera with O2 as the optical center are P1'2 and P2'2, respectively. The names of the imaging points remain unchanged after the camera imaging plane is rotated to the front of the lens (consistent with the image observed by the human eye).

[0068] When observing objects P1 and P2 from O1, you will see P1 to the right of P2; when observing P1 and P2 from O2, you will see P1 to the left of P2.

[0069] It can be seen that if the parallax at P2 is eliminated when stitching images (i.e., the positions of points P2'1 and P2'2 are made to coincide in the stitched image), the image to the right of P2'1 and the image to the left of P2'2 will be missing, that is, P1 will disappear or be incomplete in the stitched image; conversely, if the parallax at P1 is eliminated (i.e., the positions of points P1'1 and P1'2 are made to coincide in the stitched image), the image between P1'1 and P1'2 will appear on both sides of the stitching seam, that is, P2 will become two ghost images on the left and right in the stitched image.

[0070] Because images of the same object captured by cameras at different locations exhibit parallax, and objects at different distances from the lens exhibit varying parallax (larger for nearby objects and smaller for distant ones), theoretically perfect stitching of images captured by two cameras at the same moment is impossible. Eliminating parallax for nearby objects in the stitched image results in ghosting of distant objects; eliminating parallax for distant objects results in missing images of nearby objects.

[0071] The embodiments of the present application disclose an image stitching method, device, electronic device, and storage medium, which can improve the quality of image stitching. Detailed descriptions are provided below.

[0072] Figure 2a is an application scenario diagram of an image stitching method in one embodiment. As shown in Figure 2a, the multi-camera device 10 can be connected to the electronic device 20 and the server 30 for communication; the electronic device 20 can be connected to the multi-camera device 10 and the server 30 for communication.

[0073] Optionally, the server 30 and the multi-camera device 10 and the electronic device 20 can be connected to each other in communication based on a communication protocol such as HTTP / HTTPS protocol, TCP / IP protocol, WebSocket protocol, etc., which is not specifically limited.

[0074] The multi-camera device 10 and the electronic device 20 can be connected to each other through communication technologies such as wireless fidelity communication technology, Bluetooth communication technology, ZigBee communication technology, RS485 wireless transmission technology, cellular communication technology, etc., without limitation.

[0075] The multi-camera device 10 is a camera device that integrates multiple cameras, such as a binocular camera device or a trinocular camera device. The cameras in the multi-camera device 10 can work together to capture images with multiple different fields of view. In some embodiments, the cameras may include, but are not limited to, wide-angle cameras, fisheye cameras, depth cameras, etc.

[0076] The electronic device 20 may include, but is not limited to, a mobile phone, a vehicle-mounted terminal, a wearable device, a mobile station (MS), a tablet computer, a laptop computer, etc., and is not specifically limited.

[0077] The server 30 may include a physical server, a virtual server, a cloud server, etc., and is not specifically limited thereto.

[0078] In some embodiments, the multi-camera device 10 can capture a first image through a first camera and a second image through a second camera, determine a target object in the overlapping area between the first image and the second image, and determine a target feature point screening threshold for the target object based on feature information of the target object, wherein the feature information indicates the importance of the target object, and the importance of the target object is negatively correlated with the target feature point screening threshold; the multi-camera device 10 performs image stitching on the first image and the second image based on the feature points determined by the target feature point screening threshold. After the multi-camera device 10 performs image stitching on the first image and the second image, the image stitching result can be transmitted to the user's electronic device 20 for display.

[0079] In other embodiments, after the multi-camera device 10 captures a first image through a first camera and a second image through a second camera, the first image and the second image can be transmitted to the electronic device 20; the electronic device 20 can obtain the first image and the second image captured by the first camera and the second camera respectively at the same time, and determine the target object in the overlapping area between the first image and the second image, and determine the target feature point screening threshold of the target object based on the feature information of the target object, wherein the feature information indicates the importance of the target object, and the importance of the target object is negatively correlated with the target feature point screening threshold; the electronic device 20 performs image stitching on the first image and the second image based on the feature points determined by the target feature point screening threshold to obtain an image stitching result.

[0080] It should be noted that the first image taken by the first camera, the second image taken by the second camera, and the image stitching result can be stored in one or more devices among the multi-camera device 10, the electronic device 20, and the server 30, without specific limitation.

[0081] Please further refer to FIG. 2 b , which is a schematic structural diagram of a multi-camera device in one embodiment.

[0082] A multi-camera device may include an input module, a processor, a memory, and a transmission module. The input module includes a microphone and multiple cameras with the same specifications but different external parameters. The processor is used to run the algorithm and control the execution of instructions. The memory is used to store the algorithm program and image stitching results. The transmission module is used to transmit the calculation results, such as the algorithm operation results and image stitching results, to the user's mobile phone or computer.

[0083] As shown in FIG3 , in one embodiment, an image stitching method is provided, which can be applied to the above-mentioned multi-camera device or electronic device. The method may include the following steps:

[0084] Step 310: Acquire a first image and a second image captured by a first camera and a second camera respectively at the same time.

[0085] The first camera and the second camera are different cameras in a multi-camera device. The multi-camera device may include multiple cameras, and the first camera and the second camera may be any two cameras in the multi-camera device whose shooting fields of view overlap, where the overlapping field of view is the portion where the shooting fields of the first camera and the second camera overlap.

[0086] Exemplarily, a multi-camera device includes three cameras, namely camera a, camera b, and camera c. If camera b is located in the middle of the three cameras, and the fields of view of cameras a and b overlap, and the fields of view of cameras b and c overlap, then the first camera and the second camera can refer to cameras a and b, or to cameras b and c. Assume that the image captured by camera a at time t is image a, the image captured by camera b at time t is image b, and the image captured by camera c at time t is image c. The multi-camera device can separately determine the image stitching results of images a and b, and the image stitching results of images b and c, and determine a final complete image stitching result based on the multiple image stitching results determined separately.

[0087] The first image and the second image have overlapping areas. Because the fields of view of the first camera and the second camera overlap, the first image and the second image also have corresponding overlapping areas. The first image and the second image are images captured by the first camera and the second camera at the same time and from different angles of the same scene.

[0088] The internal and external parameters of the first and second cameras can be parameters determined before leaving the factory. The external parameters can be determined based on the angle and position of the camera; the internal parameters can be determined using the Zhang Zhengyou calibration method (also known as the checkerboard calibration method). The first and second cameras capture multi-angle checkerboard images. The checkerboard image is a calibration plate composed of black and white squares, and the checkerboard image is used as a calibration object for camera calibration (an object mapped from the real world to the digital image). Optionally, the first and second cameras can use the same specifications, so the internal parameters of the first and second cameras can be the same.

[0089] In some embodiments, after the step of acquiring the first image and the second image respectively captured by the first camera and the second camera at the same time, the following steps are also included: performing distortion correction on the first image and the second image based on the intrinsic parameters of the first camera and the second camera; and / or, performing perspective transformation on the first image and the second image based on the intrinsic parameters and extrinsic parameters of the first camera and the second camera; and / or, performing brightness balance on the first image and the second image.

[0090] As an optional implementation, the first image, the second image, the intrinsic parameter matrix of the first camera, and the intrinsic parameter matrix of the second camera can be input into the distortion correction function to perform radial distortion correction on the first image and the second image. Through distortion correction, the deformation of the target object in the image can be ensured to be as small as possible.

[0091] Since the bisector of the viewing angle of the first camera and the second camera is not perpendicular to the baseline, it will affect the binocular ranging and image stitching effect. As an optional implementation, the correction matrix between the first camera and the second camera can be calculated, and the first image and the second image can be perspective transformed by the correction matrix. The calculation formula of the correction matrix H is as follows: H = KRK -1 ; (1)

[0092] Since the specifications of each camera in a multi-camera system are generally the same, and the intrinsic parameter matrices of the first and second cameras are also the same, K can be the intrinsic parameter matrix of the first or second camera, and R is the rotation matrix between the first and second cameras. By performing a perspective transformation on the first and second images using the correction matrix, the first image can be mapped to the image captured when the bisector of the viewing angle of the first camera is perpendicular to the baseline, and the second image can be mapped to the image captured when the bisector of the viewing angle of the second camera is perpendicular to the baseline.

[0093] As shown in Figure 4, which is a schematic diagram of perspective transformation in one embodiment, it can be seen that the viewing angles of both cameras are facing outward, resulting in severe distortion of the target object. Through perspective transformation, the image can be mapped to the image captured when the bisector of the camera's viewing angle is perpendicular to the baseline. In this case, the camera's viewing angle is facing the target object.

[0094] The baseline is the line connecting the optical centers of the first and second cameras. The optical center is the optical center of the camera, also known as the imaging center. When a camera forms an image, light passes through the lens and is focused on the imaging plane. The pixels on the imaging plane are the image captured by the camera. The optical center of the camera is the optical center of the lens, also known as the center point of light passing through the lens.

[0095] After acquiring a first image and a second image captured by a first camera and a second camera at the same time, brightness balancing is performed on the first image and the second image. This can balance the image brightness, including balancing the brightness of each area in a single image and the brightness between the first image and the second image.

[0096] Therefore, in the embodiment of the present application, the brightness of the first image and the second image is unified, and the parallax caused by camera rotation and other problems are solved through distortion correction and perspective transformation, ensuring that the same target object is imaged consistently in different images, which helps to improve the quality of image stitching.

[0097] Step 320: Determine the target object in the overlapping area.

[0098] In some embodiments, methods for determining the target object in the overlapping area may include, but are not limited to, target detection algorithms, instance segmentation and semantic segmentation algorithms, edge detection algorithms, template matching algorithms, and the like.

[0099] The target object can be a stationary object or a moving object, and there is no specific limitation.

[0100] Step 330 : determining a target feature point screening threshold of the target object according to the feature information of the target object.

[0101] The feature information indicates the importance of the target object, and the importance of the target object is negatively correlated with the target feature point screening threshold.

[0102] The target feature point screening threshold is used to detect feature points of the target object. Feature point detection is performed on the first and second images based on the target feature point screening threshold to obtain feature points for image stitching. These feature points can reflect important areas or regions of interest in the first and second images.

[0103] Optionally, the size of the target feature point screening threshold can affect the number of feature points of the target object.

[0104] The higher the importance of the target object, the lower the target feature point screening threshold, and the more feature points of the target object are detected according to the target feature point screening threshold, which helps to mainly perform image stitching based on the feature points of the target object with high importance when stitching the first image and the second image.

[0105] In some embodiments, a highly important target object may be a focal point in an image that is desired to be highlighted, such as the foreground or subject of an image. Therefore, image stitching is performed based on the feature points of the highly important target object. This allows for a focus on the mapping relationship between the regions containing the highly important target object in the two images during image stitching. This facilitates image stitching based on important regions within the images, improves image stitching quality, and significantly enhances the stitching quality of regions containing highly important target objects.

[0106] In some embodiments, the feature information of the target object includes one or more of the following: depth information, size of the recognition frame, confidence level of the recognition frame, and object type.

[0107] The depth information may refer to the distance between the target object and a baseline, where the baseline is a line connecting the optical center of the first camera and the optical center of the second camera.

[0108] Methods for determining depth information may include, but are not limited to, parallax, time-of-flight, and structured light ranging. The parallax method involves matching corresponding pixels of the same scene in a first image and a second image through binocular matching to obtain a disparity map, which is then used to calculate depth information. The disparity map can be used to reflect the positional offset between the image feature points corresponding to the first image and the second image.

[0109] The following describes the parallax method in conjunction with FIG5 :

[0110] Assume that O1 and O2 are the optical centers of the two cameras, P is the location of the target object, the distance from P to the baseline is d, the baseline length is b, the focal length of the two cameras is f, the imaging points of P on the two camera sensors are P1 and P2 respectively (the camera imaging plane is rotated and placed in front of the lens), and the distance from P1 to P2 is m.

[0111] From the relationship between similar triangles in Figure 5, we can get:

[0112] Also: m=b-(XR-XL); (3)

[0113] We can get:

[0114] Where b and f can be obtained by camera calibration, XR and XL are the horizontal coordinates of imaging points P1 and P2 of point P in the left and right images, respectively. XR-XL can be measured from the view after binocular matching, and then the distance d from P to the baseline can be calculated according to the formula.

[0115] Since smaller depth information indicates a smaller distance between the target object and the baseline, which means the target object is closer to the multi-camera device and occupies a larger proportion in the image, in some embodiments, the smaller the depth information of the target object, the higher its importance, and the smaller the target feature point screening threshold.

[0116] In some embodiments, the feature information of the target object includes depth information, and the overlapping area includes the first target object and the second target object; determining the target feature point screening threshold of the target object based on the feature information of the target object may include: determining the first target feature point screening threshold of the first target object based on the first depth information of the first target object, and determining the second target feature point screening threshold of the second target object based on the second depth information of the second target object; the first depth information is the first distance between the first target object and the baseline, the second depth information is the second distance between the second target object and the baseline, and the baseline is the line between the optical center of the first camera and the optical center of the second camera; the first distance is less than the second distance, and the first target feature point screening threshold is less than the second target feature point screening threshold.

[0117] The first target object and the second target object may be any target objects in the overlapping area. The distance between the first target object and the baseline is the first distance, and the distance between the second target object and the baseline is the second distance. The first target feature point screening threshold is the target feature point screening threshold for the first target object, and the second target feature point screening threshold is the target feature point screening threshold for the second target object.

[0118] Since the first distance is smaller than the second distance, it means that the first target object is closer to the baseline, that is, closer to the multi-camera device. Therefore, the area of ​​the first target object in the image will be larger than the area of ​​the second target object in the image, and the importance of the first target object is higher than the importance of the second target object. Therefore, the first target feature point screening threshold is smaller than the second target feature point screening threshold. Therefore, the number of feature points of the first target object extracted according to the first target feature point screening threshold is greater than the number of feature points of the second target object extracted according to the second target feature point screening threshold. This helps to mainly perform image stitching based on the feature points of the target object closer to the multi-camera device when stitching the first image and the second image, so that the mapping relationship of the area where the target object with small depth information is located in the two images is focused on when stitching the images, which can improve the quality of image stitching and greatly improve the stitching quality of the area where the target object with small depth information is located.

[0119] The recognition box is a box generated by target detection in the overlapping area. As an optional implementation, if the target object in the overlapping area is determined by the target detection algorithm, a corresponding target candidate box (Instance Bounding Box) will be generated for the target object. It can also be called a recognition box, bounding box, or detection box. In the embodiments of this application, it can be called a recognition box. The recognition box can be used to identify the position of the target object in the overlapping area.

[0120] Since the larger the size of the recognition box is, the larger the proportion of the target object in the image is, in some embodiments, the larger the size of the recognition box is, the higher the importance is.

[0121] The confidence level of a bounding box refers to the degree of confidence in the target detection algorithm's detection results, specifically, the degree of confidence that the target object is contained within the bounding box. The confidence level is a real number between 0 and 1; the closer the confidence level is to 1, the more confidence the target detection algorithm has in the detection result, and vice versa. Target detection confidence can include classification confidence and location confidence. Classification confidence measures the accuracy of the target detection algorithm's determination of the target object's category, while location confidence measures the accuracy of the target detection algorithm's tracking of the target object's location. Therefore, a higher confidence level in a bounding box indicates a more accurate determination of the target object's category and a higher accuracy in tracking the target object's location.

[0122] Since a greater confidence level of a recognition frame indicates a higher accuracy of target object detection, in some embodiments, a greater confidence level of a recognition frame indicates a higher level of importance.

[0123] The object type of the target object may refer to the category of the target object, such as a person, an animal, a chair, a car, etc. In some embodiments, if the object type of the target object meets the preset object category, it means that the importance of the target object is higher.

[0124] In some embodiments, the characteristic information of the target object includes the object category; the multi-camera device determines the target feature point screening threshold of the target object based on the characteristic information of the target object, which may include: judging whether the object category of the target object meets the preset object category; if the object category of the target object meets the preset object category, determining the third target feature point screening threshold of the target object; if the object category of the target object does not meet the preset object category, determining the fourth target feature point screening threshold of the target object; the third target feature point screening threshold is less than the fourth target feature point screening threshold.

[0125] The preset object category can be a user-defined category of object interest. Determining the target feature point screening threshold based on the object category helps improve the stitching quality of the area containing the target object of interest. For example, in a single image, a user might be more interested in improving the stitching quality of the area containing a car. Therefore, target objects with the car category are given the highest importance. The higher the importance, the smaller the target feature point screening threshold, and the more feature points corresponding to the car are extracted. Therefore, image stitching can be focused on the car's feature points, significantly improving the stitching quality of the area containing the car.

[0126] Therefore, the third target feature point screening threshold of the target object is the target feature point screening threshold when the object category of the target object meets the preset object category; the fourth target feature point screening threshold of the target object is the target feature point screening threshold when the object category of the target object does not meet the preset object category, and the third target feature point screening threshold is less than the fourth target feature point screening threshold.

[0127] The overlapping region may include one or more target objects.

[0128] Since the area where the target object is located is more important than the background area in the image, the target feature point screening threshold of the target object is determined according to the feature information of the target object, and more feature points of the target object are screened according to the target feature point screening threshold. This helps to focus more on the feature points of the target object than the feature points of the background area when stitching the first image and the second image, thereby improving the quality of image stitching and achieving excellent stitching quality in the area where the target object is located.

[0129] If the overlapping area includes multiple target objects, the importance of different target objects is different. The higher the importance of the target object, the lower the target feature point screening threshold, and the more feature points of the target objects are detected according to the target feature point screening threshold. This helps to focus more on the feature points of target objects with high importance when stitching the first image and the second image, thereby improving the quality of image stitching and greatly improving the stitching quality of the area where the target object with the highest importance is located.

[0130] High stitching quality means that the stitching seams of the image stitching results are not obvious and the imaging of the area where the key target objects are located is clear.

[0131] Step 340 : performing image stitching on the first image and the second image based on the feature points determined by the target feature point screening threshold.

[0132] Feature points are extracted from the first image and the second image respectively according to the target feature point screening threshold, and the feature points determined in the first image are matched with the feature points determined in the second image to determine the mapping relationship between the first image and the second image, so as to perform image stitching on the first image and the second image according to the mapping relationship.

[0133] Feature point detection can identify important locations and structures in an image, helping to align and correct the images. By matching the feature points identified in the first and second images, stitching errors caused by differences in camera position, angle, or scale can be reduced. Image stitching requires calculating the mapping relationship between the two images, which is calculated based on the matching feature points in the two images.

[0134] In some embodiments, the method of determining feature points based on the target feature point screening threshold may include, but is not limited to, feature extraction methods such as the SURF (Speeded Up Robust Features) algorithm and the FAST (Features from Accelerated Segment Test) algorithm. These feature extraction methods screen out key points (corner points, edge points, etc.) or other feature points of interest in the image based on the target feature point screening threshold.

[0135] For example, the SURF algorithm can determine the feature points by the eigenvalue threshold of the Hessian Matrix. The SURF algorithm can calculate the eigenvalue of the Hessian matrix of a certain pixel in the image. The eigenvalue of the Hessian matrix can reflect the gradient change rate in the neighborhood of the pixel. The size of the eigenvalue of the Hessian matrix can be used to determine whether the pixel is a feature point. The target feature point screening threshold in the SURF algorithm refers to the eigenvalue threshold of the Hessian matrix. If the eigenvalue of the Hessian matrix is ​​large, it means that the gradient change in the neighborhood of the pixel is large, which may be a corner point or edge point; if the eigenvalue of the Hessian matrix is ​​small, it means that the gradient change in the neighborhood of the pixel is small, which may be a flat area.

[0136] If the target feature point screening threshold is set to a larger value, the number of extracted feature points will be reduced, because only pixels with very significant gradient changes in the neighborhood will be selected as feature points; if the target feature point screening threshold is set to a smaller value, the number of extracted feature points will be increased, because pixels with not very significant gradient changes in the neighborhood can also be selected as feature points.

[0137] In an embodiment of the present application, a target object in an overlapping area between a first image taken by a first camera and a second image taken by a second camera is determined, and a target feature point screening threshold of the target object is determined based on feature information of the target object, and image stitching is performed based on feature points determined by the target feature point screening threshold. Since the feature information of the target object can reflect the importance of the target object, the target feature point screening threshold of the target object is determined based on the importance of the target object, and image stitching can be focused on feature points of target objects with high importance, which helps to perform image stitching based on important areas in the image, thereby improving the quality of image stitching.

[0138] As shown in FIG6 , in one embodiment, an image stitching method is provided, which can be applied to the above-mentioned multi-camera device or electronic device. The method may include the following steps:

[0139] Step 610: Acquire a first image and a second image captured by a first camera and a second camera respectively at the same time.

[0140] The first camera and the second camera are different cameras in a multi-camera device, and there is an overlapping area between the first image and the second image.

[0141] Step 620: Determine the target object in the overlapping area.

[0142] Step 630: Determine target detection information of the target object.

[0143] In some embodiments, the target object in the overlapping area can be detected by a target detection algorithm. Optionally, the target detection algorithm may include but is not limited to a KCF (Kernel Correlation Filter) algorithm, a YOLO (You Only Look Once) algorithm, an SSD (Single Shot Multibox Detector) algorithm, and the like.

[0144] Object detection algorithms can detect and track not only static objects but also dynamic ones. By detecting objects in overlapping regions, object detection information can be obtained, including the identification box and the corresponding confidence level.

[0145] Among them, the target detection information includes a first identification box corresponding to the target object in the first image, a second identification box corresponding to the target object in the second image, a first confidence level corresponding to the first identification box, and a second confidence level corresponding to the second identification box; the first confidence level is used to indicate the degree of credibility that the target object is contained in the first identification box, and the second confidence level is used to indicate the degree of credibility that the target object is contained in the second identification box.

[0146] The first recognition frame is a recognition frame of the target object in the overlapping area of ​​the first image, and the second recognition frame is a recognition frame of the target object in the overlapping area of ​​the second image.

[0147] Since there is overlap between the first and second images, let's assume that the area in the first image that overlaps with the second image is region m, and the area in the second image that overlaps with the first image is region n. The first bounding box corresponding to the target object in the first image is the bounding box corresponding to the target object in region m; the second bounding box corresponding to the target object in the second image is the bounding box corresponding to the target object in region n.

[0148] The first confidence level is the confidence level corresponding to the first recognition frame, and the second confidence level is the confidence level corresponding to the second recognition frame.

[0149] Step 640 : Determine feature information of the target object based on the target detection information and the depth information of the target object.

[0150] The feature information indicates the importance of the target object, which is negatively correlated with the target feature point screening threshold. The feature information of the target object includes at least one of the following: depth information, recognition frame size, recognition frame confidence, and object type.

[0151] Step 650 : Determine an average area of ​​the recognition frames corresponding to the target object based on the size of the first recognition frame and the size of the second recognition frame.

[0152] The size of the first recognition frame may include the width, height, and area of ​​the first recognition frame. The size of the second recognition frame may include the width, height, and area of ​​the second recognition frame.

[0153] Determining the average area of ​​the recognition frames corresponding to the target object based on the size of the first recognition frame and the size of the second recognition frame may include: calculating the average area of ​​the first recognition frame and the area of ​​the second recognition frame to obtain the average area of ​​the recognition frames corresponding to the target object.

[0154] Step 660: Determine an average confidence level corresponding to the target object based on the first confidence level and the second confidence level.

[0155] An average value of the first confidence level and the second confidence level is calculated to obtain an average confidence level corresponding to the target object.

[0156] Step 670 : Determine a target feature point screening threshold for the target object based on the average area of ​​the recognition frame, the average confidence level, and the depth information of the target object.

[0157] In some embodiments, determining the target feature point screening threshold of the target object based on the average value of the identification box area, the average value of the confidence level, and the depth information of the target object can include: performing a weighted sum calculation on the average value of the identification box area, the average value of the confidence level, and the depth information of the target object according to the weight values ​​corresponding to the average value of the identification box area, the average value of the confidence level, and the depth information of the target object, and determining the target feature point screening threshold of the target object based on the weighted sum calculation result.

[0158] Therefore, when there are multiple target objects in the overlapping area, the average recognition box area, the average confidence level and the depth information of different target objects are different, and thus the target feature point screening thresholds of different target objects are different, and the feature points extracted for different target objects are different. Therefore, the stitching quality of important target objects can be strategically optimized during image stitching.

[0159] If the target object is dynamically moving, the average area value of the identification frame, the average confidence value, and the depth information of the target object will also change continuously. Assume that the first camera and the second camera respectively capture the first image p1 and the second image q1 at time t1. At this time, the target feature point screening threshold r1 is determined based on the average area value of the identification frame, the average confidence value, and the depth information of the target object, and the image stitching is performed based on the feature points extracted based on the target feature point screening threshold r1; assuming that the target object moves between time t1 and time t2, the average area value of the identification frame, the average confidence value, and the depth information of the target object will also change to a certain extent. Then the first camera and the second camera respectively capture the first image p2 and the second image q2 at time t2. At this time, the changed target feature point screening threshold r2 is determined based on the changed average area value of the identification frame, the average confidence value, and the depth information of the target object, and the image stitching is performed based on the feature points extracted based on the changed target feature point screening threshold r2.

[0160] Therefore, in the embodiment of the present application, the target feature point screening threshold can be updated in real time and dynamically adjusted, which helps to improve the accuracy of the extracted feature points, thereby improving the accuracy of image stitching based on the feature points.

[0161] Furthermore, in some embodiments, among the weight values ​​corresponding to the average area of ​​the identification box, the average confidence level, and the depth information of the target object, the absolute value of the weight value corresponding to the depth information is the largest. This is because depth information has a significant impact on the importance of the target object. If the target object is dynamic and moves in a direction not parallel to the baseline, the depth information of the target object will constantly change. Depth information is one of the most sensitive information that can promptly reflect the dynamic changes of the target object. Therefore, setting the weight value corresponding to the depth information of the target object to the largest value can focus more on the stitching effect of the moving target object in the overlapping area.

[0162] In some embodiments, the target feature point screening threshold of the target object is negatively correlated with the average area of ​​the recognition frame, the target feature point screening threshold of the target object is negatively correlated with the average confidence level, and the target feature point screening threshold of the target object is positively correlated with the depth information.

[0163] A larger average box area indicates a larger proportion of the target object in the image. This can be caused by, for example, a larger target object or a closer proximity to a multi-camera. A larger average box area indicates a more important target object. Therefore, a smaller target feature point screening threshold indicates a greater number of target feature points detected based on the target feature point screening threshold.

[0164] A higher average confidence value indicates a higher accuracy in target object detection, such as a higher accuracy in determining the target object's category and tracking its location. A higher average confidence value also indicates that the target object's feature points are more distinct and easier to track. Therefore, a higher average confidence value indicates a target object's importance, and therefore a smaller target feature point screening threshold indicates that more target object feature points will be detected based on the target feature point screening threshold.

[0165] The smaller the depth information, the smaller the distance between the target object and the baseline, which means that the target object is closer to the multi-camera device and the larger the proportion it occupies in the image. Therefore, the smaller the depth information of the target object, the higher its importance, the smaller the target feature point screening threshold, and the more feature points of the target object are detected according to the target feature point screening threshold.

[0166] Therefore, combining the confidence generated by the target tracking algorithm, the recognition frame size and the depth information obtained by binocular ranging calculation, different target feature point screening thresholds are set for different image areas. The feature points of target objects with high confidence, small depth information (close distance) and large area are more likely to be retained after screening.

[0167] Furthermore, the step of determining the target feature point screening threshold of the target object based on the weighted summation calculation result includes: normalizing the weighted summation calculation result; determining an adjustment coefficient based on the weighted summation calculation result after normalization, the adjustment coefficient being a positive number less than or equal to 1; determining the target feature point screening threshold of the target object based on the preset feature point screening threshold and the adjustment coefficient.

[0168] The preset feature point screening threshold can be a user-defined value, determined based on factors such as the user's requirements for the number and quality of feature points. Taking the SURF algorithm as an example, the larger the preset feature point screening threshold, the fewer feature points will be extracted, and the extracted pixels will be those with significant gradient changes in the neighborhood. The smaller the preset feature point screening threshold, the more feature points will be extracted, but it is also easy to extract some pixels with less significant gradient changes in the neighborhood as feature points.

[0169] The weighted sum calculation result is normalized so that the numerical range of the normalized weighted sum calculation result is scaled to the range of 0 to 1; an adjustment coefficient is determined based on the normalized weighted sum calculation result, and the adjustment coefficient is a positive number less than or equal to 1; and a target feature point screening threshold of the target object is determined based on the preset feature point screening threshold and the adjustment coefficient.

[0170] The adjustment coefficient is determined according to the weighted sum calculation result after normalization processing, and the preset feature point screening threshold is adjusted, so that the adjustment coefficient is determined in real time according to the average value of the recognition frame area, the average value of the confidence and the depth information of the target object, so as to dynamically adjust the preset feature point screening threshold and flexibly and accurately determine the target feature point screening threshold of the target object.

[0171] Optionally, the target feature point screening threshold of the target object is determined according to the preset feature point screening threshold and the adjustment coefficient, and the preset feature point screening threshold and the adjustment coefficient are multiplied to obtain the target feature point screening threshold of the target object.

[0172] The following is an explanation based on formulas (5) to (6).

[0173] According to the weight values ​​corresponding to the average area of ​​the recognition frame, the average confidence value, and the depth information of the target object, a weighted sum calculation result P is performed on the average area of ​​the recognition frame, the average confidence value, and the depth information of the target object. The weighted sum calculation result P can be calculated according to formula (5). P = x0 + x1S + x2C - x3d; (5)

[0174] Wherein, S is the average area of ​​the recognition box of the target object. Assuming that the width of the first recognition box corresponding to the target object in the first image is W1 and the height is H1, and the width of the second recognition box corresponding to the target object in the second image is W2 and the height is H2, then S = (W1*H1+W2*H2) / 2. C is the average confidence of the target object. Assuming that the first confidence corresponding to the target object in the first image is C1 and the second confidence corresponding to the target object in the second image is C2, then C = (C1+C2) / 2. d is the depth information of the target object, that is, the distance between the target object and the baseline. x0, x1, x2, and x3 are user-defined parameters. For example, x0, x1, x2, and x3 can be set to 1, 0.5, 0.4, and 5, respectively.

[0175] Assume that the overlapping region includes two target objects, target object A and target object B. As shown in Figure 7a, assuming that the left side shows the first image captured by the first camera, and the right side shows the second image captured by the second camera at the same time, the projections of target objects A and B on the baseline are at the same point, that is, the projections of target objects A and B in the horizontal direction parallel to the baseline are the same. As shown in Figure 7c, the left side shows the first image captured by the first camera, and the right side shows the second image captured by the second camera at the same time, the projections of target objects A and B on the baseline are at different points, that is, the projections of target objects X and Y in the horizontal direction parallel to the baseline are different. Target objects A and B can be dynamically moving target objects, moving in a direction perpendicular to the baseline or in a direction not perpendicular to the baseline.

[0176] Assume that the pixel coordinates of the upper left vertex of the recognition box of the target object A in the first image are: (X A1 , Y A1 ), the width and height of the recognition box are (W A1 , H A1 ), the confidence of the recognition box is C A1 The pixel coordinates of the upper left vertex of the recognition box of the target object B in the first image are (X B1 , Y B1 ), the width and height of the recognition box are (W B1 , H B1 ), the confidence of the recognition box is C B1 Assume that the pixel coordinates of the upper left vertex of the recognition box of the target object A in the second image are: (X A2 , Y A2 ), the width and height of the recognition box are (W A2 , H A2 ), the confidence of the recognition box is C A2 The pixel coordinates of the upper left vertex of the recognition box of the target object B in the second image are: (X B2 , Y B2 ), the width and height of the recognition box are (W B2 , H B2 ), the confidence of the recognition box is C B2 .

[0177] Combining Figure 5 and formulas (2) to (4), the depth information of the target object A can be calculated as d A , the depth information of the target object B is d B .

[0178] According to formula (5), the weighted sum calculation result P corresponding to the target object A is determined A And the weighted sum calculation result P of the target object B BAfter that, the normalized exponential function (softmax function) is used to normalize the weighted summation result P A And the weighted sum calculation result P B , as shown in formula (6).

[0179] Among them, softmax(P i ) is the normalized weighted summation result corresponding to the target object i, i = A or B. e is a natural constant, its value is approximately equal to 2.71828; when i = A, P i =P A When i=B, P i =P B ;∑ i e Pi for and sum.

[0180] The target feature point screening threshold can be calculated according to formula (7). i =Q0*(1-softmax(P i )); (7)

[0181] Among them, Q i Filter threshold for target feature points of target object i, i = A or B; (1-softmax(P i )) is the adjustment coefficient; Q0 is the preset feature point screening threshold, which can be 3000 or 3500. It can be seen that the larger the normalized weighted sum calculation result softmax(P), the smaller the target feature point screening threshold, which means that more feature points can be identified.

[0182] Therefore, based on the target feature point screening threshold for target object A, the feature points corresponding to target object A are extracted; based on the target feature point screening threshold for target object B, the feature points corresponding to target object B are extracted. Assuming that the target feature point screening threshold for target object A is lower than that for target object B, more feature points corresponding to target object A are extracted than those corresponding to target object B, allowing for greater emphasis on image stitching based on the feature points corresponding to target object A. Therefore, by determining different target feature point screening thresholds for different target objects and thereby extracting different feature points, the stitching quality of important targets can be strategically optimized.

[0183] Step 680 : performing image stitching on the first image and the second image based on the feature points determined by the target feature point screening threshold.

[0184] Therefore, the embodiment of the present application combines target tracking and binocular ranging to update the target feature point screening threshold in real time, and dynamically adjusts the target feature point screening threshold as the depth information, the average value of the recognition frame area, and the average value of the confidence level change; and, assigns a different target feature point screening threshold to each target object, so that when stitching images, more attention is paid to the image stitching quality of the area where the important target object is located, while also balancing the image stitching quality of other areas.

[0185] FIG7 b shows the image stitching result after the left and right images in FIG7 a are stitched together, and FIG7 d shows the image stitching result after the left and right images in FIG7 c are stitched together.

[0186] As shown in FIG8 , in one embodiment, an image stitching method is provided, which can be applied to the above-mentioned multi-camera device or electronic device. The method may include the following steps:

[0187] Step 801: Acquire a first image and a second image captured by a first camera and a second camera respectively at the same time.

[0188] The first camera and the second camera are different cameras in a multi-camera device, and there is an overlapping area between the first image and the second image.

[0189] Step 802: Determine the target object in the overlapping area.

[0190] Step 803: Determine target detection information of the target object.

[0191] Among them, the target detection information includes a first identification box corresponding to the target object in the first image, a second identification box corresponding to the target object in the second image, a first confidence level corresponding to the first identification box, and a second confidence level corresponding to the second identification box; the first confidence level is used to indicate the degree of credibility that the target object is contained in the first identification box, and the second confidence level is used to indicate the degree of credibility that the target object is contained in the second identification box.

[0192] Step 804 : determining feature information of the target object according to the target detection information and the depth information of the target object.

[0193] The feature information indicates the importance of the target object, which is negatively correlated with the target feature point screening threshold. The feature information of the target object includes at least one of the following: depth information, recognition frame size, recognition frame confidence, and object type.

[0194] Step 805 : determining a target feature point screening threshold of the target object according to the feature information of the target object.

[0195] The specific implementation of steps 801 to 805 can refer to the above embodiment and will not be described in detail.

[0196] Step 806 : extracting feature points within the first recognition frame in the first image according to the target feature point screening threshold, and extracting feature points outside the first recognition frame in the first image according to the preset feature point screening threshold.

[0197] The description of the preset feature point screening threshold can be referred to the above embodiment.

[0198] According to formula (7), the target feature point screening threshold of the target object is determined according to the preset feature point screening threshold and the adjustment coefficient. The adjustment coefficient is a positive number less than or equal to 1, so the target feature point screening threshold is less than or equal to the preset feature point screening threshold.

[0199] Therefore, more feature points are extracted from the target object area than from the background area, so that the stitching quality of the important target object area is emphasized while the stitching quality of the background area is also considered, thereby comprehensively improving the quality of image stitching.

[0200] Step 807 : Determine the feature points within the first recognition frame and the feature points outside the first recognition frame as first feature points corresponding to the target object in the first image.

[0201] Step 808 : extracting feature points within the second recognition frame in the second image according to the target feature point screening threshold, and extracting feature points outside the second recognition frame in the second image according to the preset feature point screening threshold.

[0202] Step 809 : Determine the feature points within the second identification frame and the feature points outside the second identification frame as second feature points corresponding to the target object in the second image.

[0203] Therefore, in the embodiment of the present application, feature points within each identification box can be detected based on the target feature point screening threshold, and feature points outside the identification box in the image can be detected based on the preset feature point screening threshold. The area within the identification box in the image can be considered as the foreground area or subject area of ​​the image, and the area outside the identification box in the image can be considered as the background area of ​​the image.

[0204] For illustration, using the application scenarios of Figures 7a and 7c, assume that the overlapping region includes two target objects, namely target object A and target object B. Therefore, based on the target feature point screening threshold for target object A, the feature points of target object A within the corresponding identification box of each image are extracted; based on the target feature point screening threshold for target object B, the feature points of target object B within the corresponding identification box of each image are extracted; based on the preset feature point screening threshold, the feature points of target object A outside the corresponding identification box of each image are extracted; and based on the preset feature point screening threshold, the feature points of target object B outside the corresponding identification box of each image are extracted.

[0205] In Figure 7a, because target object A is farther away, has a smaller area, and has a lower tracking confidence score, its feature points are more likely to be filtered than those of target object B. Therefore, when performing image stitching based on feature points, the total number of feature points occupied by target object A is less than that occupied by target object B. Therefore, the mapping relationship between the region containing target object B in the two images is given greater consideration during image stitching. However, since target object A has more matching feature points than the background area, the final stitched image prioritizes the stitching quality of the regions containing target objects A and B, while also considering the stitching quality of the background area, thereby comprehensively improving the quality of image stitching. Furthermore, since the projections of target objects A and B on the baseline are at the same point, the image stitching can easily result in missing or ghosting of certain target objects. However, the embodiments of the present application prioritize the mapping relationship between the regions containing more important target objects, such as target object B, in the two images during image stitching, effectively resolving the problem of missing or ghosting of certain target objects that may occur during image stitching.

[0206] In Figure 7c, because the projections of target objects A and B on the baseline are at different points, and their relative positions in the left and right images are the same, with target object B located to the left of target object A, the likelihood of missing or ghosting a target object is lower during image stitching. Compared to Figure 7a, target object A has a larger area, so the target feature point screening threshold for target object A in Figure 7c is lower than that in Figure 7a. Consequently, more feature points are detected in Figure 7c than in Figure 7a. Therefore, when performing image stitching in Figure 7c, the image mapping relationship of the area containing target object A is considered more than in Figure 7a.

[0207] Therefore, the embodiment of the present application can detect feature points within each identification box based on the target feature point screening threshold, and detect feature points outside the identification box in the image based on the preset feature point screening threshold. The feature points extracted from the area where the target object is located are more than the feature points extracted from the background area, thereby focusing on solving the stitching quality of the important target object area, and taking into account the stitching quality of the background area, thereby comprehensively improving the quality of image stitching.

[0208] Step 810 : Matching first feature points corresponding to the target object in the first image with second feature points corresponding to the target object in the second image, so as to perform image stitching on the first image and the second image.

[0209] In some embodiments, matching a first feature point corresponding to a target object in a first image with a second feature point corresponding to the target object in a second image to perform image stitching on the first image and the second image may include: matching the first feature point corresponding to the target object in the first image with the second feature point corresponding to the target object in the second image to obtain a homography matrix of the second image relative to the first image; performing a perspective transformation on the second image according to the homography matrix to map the second image to a camera coordinate system corresponding to the first image; and aligning pixel points corresponding to the same coordinate values ​​in the first image and the transformed second image to perform image stitching on the first image and the second image.

[0210] The homography matrix describes the transformation relationship (mapping relationship) of coplanar points from an image taken from one perspective to an image taken from another perspective.

[0211] In some embodiments, after determining the first feature point corresponding to the target object in the first image and the second feature point corresponding to the target object in the second image, a FLANN (fast library for approximate nearest neighbors) algorithm can be used to find mutually matching feature point pairs between the first feature point and the second feature point, and then the homography matrix of the second image relative to the first image can be calculated based on the feature point pairs.

[0212] Methods for calculating the homography matrix may include, but are not limited to, a RANSAC (random sample consensus) algorithm, a DLT (direct linear transform), and the like.

[0213] The second image is perspective transformed by the homography matrix so that the second image after perspective transformation is in the same camera coordinate system as the first image, and the pixels of the first image are copied to the second image after perspective transformation, thereby realizing image stitching and obtaining the image stitching result.

[0214] By calculating the homography matrix of the entire second image relative to the first image and using only the same homography matrix for transformation, the seam distribution is simple, which facilitates seam optimization and makes the seams of image stitching transition smoothly, thus solving the problem of obvious seams and improving the image stitching effect.

[0215] In some embodiments, after the step of stitching the first image and the second image, the following steps are also included: determining the first pixel value of the pixel point of the target column in the first image and the second pixel value of the pixel point of the target column in the alignment area of ​​the first image and the second image, the alignment area being the area where the pixels in the first image and the second image are aligned; the pixel point of the target column is any column of pixel points in the alignment area; performing weighted fusion on the pixel points of the target column according to the first weight and the first pixel value of the pixel point of the target column, and according to the second weight and the second pixel value of the pixel point of the target column to obtain the target pixel value corresponding to the pixel point of the target column.

[0216] In some embodiments, before the step of weightedly fusing the pixel points of the target column according to the first weight and the first pixel value of the pixel points of the target column, and according to the second weight and the second pixel value of the pixel points of the target column, the following steps are also included: determining the channel score according to the difference in the number of columns between the target column and the leftmost column in the alignment area, and the difference in the number of columns between the leftmost column and the rightmost column in the alignment area; determining the first weight and the second weight according to the channel score.

[0217] Therefore, the pixels of the non-aligned areas of the first and second images can be copied to the corresponding stitching area to fill the stitching image, and the pixel values ​​in the aligned area are adjusted to the weighted fusion result of the corresponding areas in the first and second images.

[0218] Assuming that the first image is on the left and the second image is on the right, for the aligned area in the image stitching result, in order to achieve the goal of the pixel weight of the first image being higher the closer to the left, and vice versa, the pixel weight of the second image being higher.

[0219] The target column is any column in the alignment area.

[0220] The horizontal coordinate of the pixel of the target column in the image stitching result is i, the first pixel value of the pixel point of the target column in the first image is pixel_l, and the second pixel value of the pixel point of the target column in the second image is pixel_r.

[0221] The target pixel value pixel_t corresponding to the pixel point of the target column is shown in formula (8). pixel_t=(1-alpha_i)*pixel_l+alpha_i*pixel_r; (8)

[0222] Among them, 1-alpha_i is the first weight, alpha_i is the second weight, and the first weight and the second weight are determined according to the channel score alpha_i.

[0223] The calculation formula of the channel score alpha_i is shown in (9). alpha_i=(ia) / (ba); (9)

[0224] Among them, alpha_i is the channel score, i is the horizontal coordinate of the pixel of the target column in the image stitching result, a is the horizontal coordinate of the pixel of the leftmost column in the alignment area in the image stitching result, (ia) is the difference in the number of columns between the target column and the leftmost column in the alignment area, b is the horizontal coordinate of the pixel of the leftmost column in the alignment area in the image stitching result, and (ba) is the difference in the number of columns between the leftmost column and the rightmost column in the alignment area.

[0225] For example, if the horizontal coordinate of the pixel of the target column in the image stitching result is a, that is, the pixel of the 0th column in the alignment area, then alpha_i=0, and the target pixel value corresponding to the pixel of the target column is pixel_t=1*pixel_l+0*pixel_r.

[0226] Therefore, weighted fusion processing is performed on the aligned areas in the image stitching result to make the transition of image stitching smoother.

[0227] The image stitching method of the embodiment of the present application is described below with reference to FIG9 .

[0228] Multiple cameras can be configured on a multi-camera device. Figure 9 illustrates this using two cameras as an example. After the camera collects the data, it performs operations such as distortion correction, brightness balance, and rotation correction to optimize image quality. A ranging algorithm is then used to obtain the distance from the moving target in the overlapped area to the camera device. Combined with target tracking, the target's identification frame in the image, confidence level, and other factors are used to calculate the feature point screening threshold. Because different targets have different feature point screening thresholds, a different number of feature points are detected for each target during feature point detection. The resulting calculated single-image matrix will take into account the perspective transformation relationship of areas with more feature points, enabling dynamic adjustment of the stitching distance (the distance from the moving target to the camera device), resulting in excellent imaging of significant moving targets in the overlapped area.

[0229] As shown in Figure 10, in one embodiment, an image stitching device 1000 is provided, which can be applied to the above-mentioned multi-camera device or electronic device. The image stitching device 1000 may include an image acquisition module 1010, an object determination module 1020, a threshold determination module 1030 and an image stitching module 1040.

[0230] An image acquisition module 1010 is configured to acquire a first image and a second image captured by a first camera and a second camera at the same time, respectively, where the first camera and the second camera are different cameras in a multi-camera device, and the first image and the second image have overlapping areas;

[0231] An object determination module 1020 is configured to determine a target object in the overlapped area;

[0232] A threshold determination module 1030 is configured to determine a target feature point screening threshold of the target object based on feature information of the target object, wherein the feature information indicates the importance of the target object, and the importance of the target object is negatively correlated with the target feature point screening threshold;

[0233] The image stitching module 1040 is configured to stitch the first image and the second image together based on the feature points determined by the target feature point screening threshold.

[0234] In one embodiment, the characteristic information of the target object includes one or more of the following:

[0235] Depth information, size of the recognition box, confidence of the recognition box and object type.

[0236] In one embodiment, the feature information of the target object includes depth information, and the overlapping area includes the first target object and the second target object; the threshold determination module 1030 is also used to: determine the first target feature point screening threshold of the first target object based on the first depth information of the first target object, and determine the second target feature point screening threshold of the second target object based on the second depth information of the second target object; the first depth information is the first distance between the first target object and the baseline, the second depth information is the second distance between the second target object and the baseline, and the baseline is the line between the optical center of the first camera and the optical center of the second camera; the first distance is less than the second distance, and the first target feature point screening threshold is less than the second target feature point screening threshold.

[0237] In one embodiment, the characteristic information of the target object includes at least one of the following: depth information, the size of the identification frame, the confidence of the identification frame and the object type; the threshold determination module 1030 is also used to determine the target detection information of the target object; the target detection information includes a first identification frame corresponding to the target object in the first image, a second identification frame corresponding to the target object in the second image, a first confidence corresponding to the first identification frame and a second confidence corresponding to the second identification frame; the first confidence is used to indicate the degree of confidence that the target object is contained in the first identification frame, and the second confidence is used to indicate the degree of confidence that the target object is contained in the second identification frame; the characteristic information of the target object is determined based on the target detection information of the target object and the depth information of the target object.

[0238] In one embodiment, the threshold determination module 1030 is further used to determine the average area of ​​the identification frame corresponding to the target object based on the size of the first identification frame and the size of the second identification frame; determine the average confidence level corresponding to the target object based on the first confidence level and the second confidence level; and determine the target feature point screening threshold of the target object based on the average area of ​​the identification frame, the average confidence level and the depth information of the target object.

[0239] In one embodiment, the target feature point screening threshold of the target object is negatively correlated with the average area of ​​the recognition frame, the target feature point screening threshold of the target object is negatively correlated with the average confidence level, and the target feature point screening threshold of the target object is positively correlated with the depth information.

[0240] In one embodiment, the threshold determination module 1030 is further used to perform weighted summation calculation on the average area of ​​the identification box, the average confidence value and the depth information of the target object according to the weight values ​​corresponding to the average area of ​​the identification box, the average confidence value and the depth information of the target object, and determine the target feature point screening threshold of the target object according to the weighted summation calculation result; wherein, among the weight values ​​corresponding to the average area of ​​the identification box, the average confidence value and the depth information of the target object, the absolute value of the weight value corresponding to the depth information is the largest.

[0241] In one embodiment, the threshold determination module 1030 is also used to normalize the weighted sum calculation result; determine the adjustment coefficient based on the weighted sum calculation result after normalization, and the adjustment coefficient is a positive number less than or equal to 1; determine the target feature point screening threshold of the target object based on the preset feature point screening threshold and the adjustment coefficient.

[0242] In one embodiment, the image stitching device 1000 further includes: a feature extraction module;

[0243] A feature extraction module is used to extract feature points located within a first identification frame in a first image based on a target feature point screening threshold, and to extract feature points located outside the first identification frame in the first image based on a preset feature point screening threshold; determine the feature points within the first identification frame and the feature points outside the first identification frame as first feature points corresponding to the target object in the first image; extract feature points located within a second identification frame in a second image based on the target feature point screening threshold, and to extract feature points located outside the second identification frame in the second image based on a preset feature point screening threshold; determine the feature points within the second identification frame and the feature points outside the second identification frame as second feature points corresponding to the target object in the second image.

[0244] In some embodiments, the image stitching module 1040 is further configured to match first feature points corresponding to the target object in the first image with second feature points corresponding to the target object in the second image, so as to stitch the first image and the second image.

[0245] In some embodiments, the image stitching module 1040 is further used to match the first feature point corresponding to the target object in the first image with the second feature point corresponding to the target object in the second image to obtain the homography matrix of the second image relative to the first image; perform perspective transformation on the second image according to the homography matrix to map the second image to the camera coordinate system corresponding to the first image; align the pixel points corresponding to the same coordinate values ​​in the first image and the transformed second image to perform image stitching on the first image and the second image.

[0246] In some embodiments, the image stitching apparatus 1000 further includes: a weighted fusion module;

[0247] A weighted fusion module is used to determine the first pixel value of the pixel point of the target column in the first image and the second pixel value of the pixel point of the target column in the aligned area of ​​the first image and the second image, where the aligned area is the area where the pixels in the first image and the second image are aligned; the pixel point of the target column is any column of pixel points in the aligned area; the pixel points of the target column are weightedly fused according to the first weight and the first pixel value of the pixel point of the target column, and according to the second weight and the second pixel value of the pixel point of the target column to obtain the target pixel value corresponding to the pixel point of the target column.

[0248] In some embodiments, the weighted fusion module is further used to determine the channel score based on the difference in the number of columns between the target column and the leftmost column in the alignment area, and the difference in the number of columns between the leftmost column and the rightmost column in the alignment area; and determine the first weight and the second weight based on the channel score.

[0249] In some embodiments, the characteristic information of the target object includes an object category; the threshold determination module 1030 is also used to determine whether the object category of the target object meets the preset object category; if the object category of the target object meets the preset object category, the third target feature point screening threshold of the target object is determined; if the object category of the target object does not meet the preset object category, the fourth target feature point screening threshold of the target object is determined; the third target feature point screening threshold is less than the fourth target feature point screening threshold.

[0250] In some embodiments, the image acquisition module 1010 is also used to perform distortion correction on the first image and the second image based on the internal parameters of the first camera and the second camera; and / or, perform perspective transformation on the first image and the second image based on the internal parameters and external parameters of the first camera and the second camera; and / or, perform brightness balance on the first image and the second image.

[0251] In an embodiment of the present application, a target object in an overlapping area between a first image taken by a first camera and a second image taken by a second camera is determined, and a target feature point screening threshold of the target object is determined based on feature information of the target object, and image stitching is performed based on feature points determined by the target feature point screening threshold. Since the feature information of the target object can reflect the importance of the target object, the target feature point screening threshold of the target object is determined based on the importance of the target object, and image stitching can focus on feature points of target objects with high importance, which helps to perform image stitching based on important areas in the image and can improve the quality of image stitching.

[0252] Please refer to Figure 11, which is a block diagram of the structure of an electronic device according to an embodiment. As shown in Figure 11, the electronic device 20 may include: a memory 1110 storing executable program code; and a processor 1120 coupled to the memory 1110. The processor 1120 invokes the executable program code stored in the memory 1110 to execute any of the image stitching methods disclosed in the embodiments of this application.

[0253] Please refer to Figure 12, which is a block diagram of the structure of a multi-camera device in one embodiment. As shown in Figure 12, the multi-camera device 10 may include a memory 1210, a processor 1220, and a multi-camera 1230. The multi-camera 1230 is configured to capture multiple images containing overlapping areas. The memory 1210 stores computer-readable instructions that, when executed by one or more processors 1220, cause the one or more processors 1220 to execute any of the image stitching methods disclosed in the embodiments of this application to stitch multiple images together.

[0254] An embodiment of the present application discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by the processor, the processor implements any one of the image stitching methods disclosed in the embodiment of the present application.

[0255] It should be understood that the references to "one embodiment" or "an embodiment" throughout the specification mean that the specific features, structures, or characteristics associated with the embodiment are included in at least one embodiment of the present application. Therefore, the references to "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. Those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required for the present application.

[0256] In the various embodiments of the present application, it should be understood that the size of the sequence numbers of the above-mentioned processes does not necessarily mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, the functional units in the various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0257] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a memory and includes several requests for a computer device (which can be a personal computer, a multi-camera device or a network device, etc., specifically a processor in a computer device) to execute some or all of the steps of the above-mentioned methods of the various embodiments of the present application.

[0258] The above describes in detail an image stitching method, device, electronic device, and storage medium disclosed in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is intended only to help understand the method and core concept of the present application. At the same time, for those skilled in the art, based on the concept of the present application, there may be changes in the specific implementation methods and scope of application. In summary, the contents of this specification should not be construed as limiting the present application.

Claims

1. An image stitching method, characterized in that, The method includes: Obtaining a first image and a second image respectively captured by a first camera and a second camera at the same moment, where the first camera and the second camera are different cameras in a multi-camera device, and there is an overlapping area between the first image and the second image; Determining a target object in the overlapping area; Determining a target feature point screening threshold for the target object according to the feature information of the target object, where the feature information indicates the importance level of the target object, and the importance level of the target object is negatively correlated with the target feature point screening threshold; Performing image stitching on the first image and the second image according to the feature points determined by the target feature point screening threshold.

2. The method according to claim 1, wherein The feature information of the target object includes one or more of the following: Depth information, size of the recognition frame, confidence of the recognition frame, and object type.

3. The method according to claim 1, wherein The feature information of the target object includes depth information, and the overlapping area includes a first target object and a second target object; the determining the target feature point screening threshold for the target object according to the feature information of the target object includes: Determining a first target feature point screening threshold for the first target object according to the first depth information of the first target object, and determining a second target feature point screening threshold for the second target object according to the second depth information of the second target object; the first depth information is the first distance between the first target object and the baseline, the second depth information is the second distance between the second target object and the baseline, the baseline is the line connecting the optical center of the first camera and the optical center of the second camera; the first distance is less than the second distance, and the first target feature point screening threshold is less than the second target feature point screening threshold.

4. The method according to any one of claims 1-3, characterized in that, The feature information of the target object includes at least one of the following: depth information, size of the recognition frame, confidence of the recognition frame, and object type. Before determining the feature point screening threshold for the target object according to the feature information of the target object, the method further includes: Determining the target detection information of the target object; the target detection information includes a first recognition frame corresponding to the target object in the first image, a second recognition frame corresponding to the target object in the second image, a first confidence corresponding to the first recognition frame, and a second confidence corresponding to the second recognition frame; the first confidence is used to indicate the credibility of including the target object in the first recognition frame, and the second confidence is used to indicate the credibility of including the target object in the second recognition frame; Determining the feature information of the target object according to the target detection information of the target object and the depth information of the target object.

5. The method according to claim 4, wherein The determining the target feature point screening threshold for the target object according to the feature information of the target object includes: Determining the average area of the recognition frame corresponding to the target object according to the size of the first recognition frame and the size of the second recognition frame; Determining the average confidence corresponding to the target object according to the first confidence and the second confidence; Determine a target feature point screening threshold for the target object according to the average value of the recognition frame area, the average value of the confidence level, and the depth information of the target object.

6. The method according to claim 5, wherein The target feature point screening threshold for the target object is negatively correlated with the average value of the recognition frame area, the target feature point screening threshold for the target object is negatively correlated with the average value of the confidence level, and the target feature point screening threshold for the target object is positively correlated with the depth information.

7. The method according to claim 5, characterized in that, The determining the target feature point screening threshold for the target object according to the average value of the recognition frame area, the average value of the confidence level, and the depth information of the target object includes: Perform a weighted sum calculation on the average value of the recognition frame area, the average value of the confidence level, and the depth information of the target object according to the weight values corresponding to the average value of the recognition frame area, the average value of the confidence level, and the depth information of the target object respectively, and determine the target feature point screening threshold for the target object according to the result of the weighted sum calculation; wherein, among the weight values corresponding to the average value of the recognition frame area, the average value of the confidence level, and the depth information of the target object respectively, the absolute value of the weight value corresponding to the depth information is the largest.

8. The method according to claim 7, wherein The determining the target feature point screening threshold for the target object according to the result of the weighted sum calculation includes: Perform a normalization process on the result of the weighted sum calculation; Determine an adjustment coefficient according to the result of the weighted sum calculation after the normalization process, and the adjustment coefficient is a positive number less than or equal to 1; Determine the target feature point screening threshold for the target object according to a preset feature point screening threshold and the adjustment coefficient.

9. The method according to claim 4, wherein Before performing image stitching on the first image and the second image according to the feature points determined by the target feature point screening threshold, the method further includes: Extract the feature points within the first recognition frame in the first image according to the target feature point screening threshold, and Extract the feature points outside the first recognition frame in the first image according to a preset feature point screening threshold; Determine the feature points within the first recognition frame and the feature points outside the first recognition frame as the first feature points corresponding to the target object in the first image; Extract the feature points within the second recognition frame in the second image according to the target feature point screening threshold, and extract the feature points outside the second recognition frame in the second image according to the preset feature point screening threshold; Determine the feature points within the second recognition frame and the feature points outside the second recognition frame as the second feature points corresponding to the target object in the second image.

10. An image stitching device, characterized in that, The apparatus includes: An image acquisition module, configured to acquire a first image and a second image respectively captured by a first camera and a second camera at the same time, the first camera and the second camera are different cameras in a multi-camera device, and there is an overlapping area between the first image and the second image; An object determination module, configured to determine a target object in the overlapping area; A threshold determination module, configured to determine a target feature point screening threshold of the target object according to the feature information of the target object, where the feature information indicates the importance level of the target object, and the importance level of the target object is negatively correlated with the target feature point screening threshold; An image stitching module, configured to perform image stitching on the first image and the second image according to the feature points determined by the target feature point screening threshold.

11. An electronic device, characterized in that, Comprising a memory and one or more processors, wherein computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 9.

12. A multi-camera device, characterized in that, Comprising a memory, a processor, and a multi-camera: The multi-camera is configured to capture multiple images including an overlapping area, and computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 9 to implement stitching of the multiple images.

13. One or more non-transitory computer-readable storage media storing computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the method according to any one of claims 1 to 9.