Intelligent multi-camera tracheal intubation image processing method, device and system

By acquiring and combining images from multiple cameras for stereo matching and depth calculation, a 3D reconstruction map is generated and the target is identified. This solves the problems of inaccurate depth perception and insufficient real-time performance of multi-camera systems during endotracheal intubation, achieving high-precision and efficient image processing results.

CN120833436APending Publication Date: 2025-10-24TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510681627.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Multi-view cameras suffer from inaccurate depth perception and insufficient real-time performance in endotracheal intubation image processing, which limits their performance and practicality in complex application scenarios.

Method used

Multiple images are acquired and combined by a multi-view image acquisition unit, stereo matching and depth calculation are performed to generate a depth map, and 3D reconstruction is performed by combining the depth map and the original image. Target recognition is then performed within the 3D reconstructed image group, and the target depth is output.

Benefits of technology

This improves the depth perception accuracy and real-time processing capability of multi-view cameras in endotracheal intubation image processing, ensuring the accuracy and efficiency of endotracheal intubation operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833436A_ABST
    Figure CN120833436A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent multi-view camera tracheal intubation image processing method, device and system, and relates to the technical field of medical instruments, the method comprises the following steps: in a tracheal intubation process, a multi-view image acquisition unit collects a plurality of images and combines the images to form an image pair; and carrying out stereo matching and depth calculation on the image pairs to generate a depth map. And performing three-dimensional reconstruction by combining the depth map and the original image to generate a three-dimensional model. And finally, identifying the target in the three-dimensional model and outputting the depth of the target as an image processing result. The technical problems that an existing multi-view camera is inaccurate in depth perception and insufficient in real-time performance during tracheal intubation image processing are solved, and the technical effect of improving the depth perception precision and the real-time processing capacity of the multi-view camera during tracheal intubation image processing is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical devices, in particular to an intelligent multi-camera endotracheal tube image processing method, device and system. BACKGROUND

[0002] In the field of image processing, multi-camera systems provide rich spatial information by simultaneously capturing images from multiple perspectives, and are widely used in tasks such as three-dimensional reconstruction, depth perception and target recognition. Compared with monocular camera systems, multi-camera systems can effectively solve the problem of insufficient depth information, significantly improving the accuracy and reliability of image processing. However, multi-camera systems still face many technical challenges in practical applications, including the need for efficient stereo matching algorithms, the dependence of accurate depth calculation on accurate camera parameter calibration, and how to effectively fuse multi-perspective depth information in the three-dimensional reconstruction process to reduce data redundancy and noise interference. These problems limit the performance and practicality of multi-camera systems in complex application scenarios, and optimized image processing methods and efficient algorithms are needed to improve their overall performance.

[0003] In the related art at this stage, there are technical problems of inaccurate depth perception and insufficient real-time performance of multi-camera in endotracheal tube image processing. SUMMARY

[0004] The present application provides an intelligent multi-camera endotracheal tube image processing method, device and system, which solves the technical problems of inaccurate depth perception and insufficient real-time performance of existing multi-camera in endotracheal tube image processing.

[0005] The present application provides an intelligent multi-camera endotracheal tube image processing method, device and system, which solves the technical problems of inaccurate depth perception and insufficient real-time performance of existing multi-camera in endotracheal tube image processing.

[0006] In endotracheal intubation, the multi-image acquisition unit acquires multiple images and combines to obtain at least one image combination; the at least one image combination is subjected to stereo matching processing and depth calculation to obtain at least one depth map; the at least one depth map and multiple images are subjected to grayscale processing to obtain at least one disparity map, three-dimensional reconstruction is performed to obtain a three-dimensional reconstruction map set; target recognition is performed in the three-dimensional reconstruction map set, and the target depth obtained is output as an image processing result.

[0007] The present application provides an intelligent multi-camera endotracheal tube image processing device, which comprises a multi-image acquisition unit and further comprises:

[0008] The image combination acquisition module is configured to acquire a plurality of images by the multi-view image acquisition unit in the tracheal intubation, and combine to obtain at least one image combination; the depth map acquisition module is configured to perform stereo matching processing and depth calculation on the at least one image combination, and obtain at least one depth map; the three-dimensional reconstruction map set acquisition module is configured to perform grayscale processing on the at least one depth map and the plurality of images, obtain at least one disparity map, perform three-dimensional reconstruction, and obtain a three-dimensional reconstruction map set; and the target depth output module is configured to perform target identification in the three-dimensional reconstruction map set, and output a target depth obtained as an image processing result.

[0009] The present application provides an intelligent multi-view camera tracheal intubation image processing system, comprising:

[0010] The memory is configured to store executable instructions, and the processor is configured to execute the executable instructions stored in the memory to implement the intelligent multi-view camera tracheal intubation image processing method.

[0011] The intelligent multi-view camera tracheal intubation image processing method, device and system provided by the present application first acquire a plurality of images by the multi-view image acquisition unit during the tracheal intubation process, and combine the images to form image pairs. The image pairs are subjected to stereo matching and depth calculation to generate depth maps. Three-dimensional models are generated by combining the depth maps and the original images for three-dimensional reconstruction. Finally, targets are identified in the three-dimensional models, and depths of the targets are output as image processing results, thereby achieving the technical effects of improving the depth perception accuracy and real-time processing capability of the multi-view camera in the tracheal intubation image processing. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments of the present disclosure will be briefly introduced below. In the present application, a flowchart is used to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the foregoing or the following operations are not necessarily performed in sequence. On the contrary, various steps can be processed in reverse order or simultaneously as needed. Meanwhile, other operations can be added to these processes, or one or more steps can be removed from these processes.

[0013] Figure 1 A flowchart of an intelligent multi-view camera tracheal intubation image processing method provided by an embodiment of the present application is shown in the figure;

[0014] Figure 2 A structural diagram of an intelligent multi-view camera tracheal intubation image processing device provided by an embodiment of the present application is shown in the figure;

[0015] Figure 3The structure diagram of the intelligent multi-view camera tracheal intubation image processing system provided by the embodiment of the application is shown.

[0016] Label explanation: image combination acquisition module 10, depth map acquisition module 20, three-dimensional reconstruction map set acquisition module 30, target depth output module 40, input device 301, processor 302, memory 303, output device 304. DETAILED DESCRIPTION

[0017] The above description is only a summary of the technical scheme of the application, in order to more clearly understand the technical means of the application, and can be implemented according to the content of the specification, and in order to make the above and other purposes, characteristics and advantages of the application more obvious and easy to understand, the following specific embodiments of the application are described.

[0018] In order to make the purpose, technical scheme and advantages of the application more clear, the following will combine the drawings to further describe the application in detail, the described embodiments should not be regarded as the limitation of the application, all other embodiments obtained by the person skilled in the art without creative labor belong to the protection scope of the application.

[0019] In the following description, "some embodiments" are related to a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same subset or different subset of all possible embodiments, and can be combined with each other without conflict, the term "first\second" involved only distinguishes similar objects, and does not represent the specific order of the object. The terms "include" and "have" and any variants, are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server containing a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the application belongs. The terms used herein are only for the purpose of describing the embodiments of the application.

[0020] The embodiment of the application provides a smart multi-view camera tracheal intubation image processing method, as shown in the figure, the method comprises: Figure 1 The method comprises:

[0021] Step S100, in the tracheal intubation, a plurality of images are acquired by the multi-view image acquisition unit, and at least one image combination is obtained by combination. Specifically, during the tracheal intubation process, the multi-view image acquisition unit synchronously acquires a plurality of images of the target region through a plurality of cameras (such as two or three), each camera captures data from a different perspective to ensure the coverage of the spatial information of the target region. After the acquisition is completed, the system combines these images two by two to generate image combinations, for example, for three cameras A, B, and C, combinations of A-B, B-C, and A-C can be formed. These image combinations provide input basis for subsequent stereo matching and depth calculation, generate depth maps through disparity calculation, and provide rich data support for three-dimensional reconstruction and target recognition. The application of multi-view cameras and image combinations solves the problem of limited perspective of monocular cameras, and improves the integrity of spatial information and the accuracy of subsequent processing.

[0022] In one possible implementation, in the tracheal intubation, a plurality of images are acquired by the multi-view image acquisition unit, and at least one image combination is obtained by combination, step S100 further includes step S110, a plurality of images are acquired by the multi-view image acquisition unit. Specifically, during the process of acquiring a plurality of images by the multi-view image acquisition unit, a plurality of cameras are used to synchronously capture the target region from different positions and angles, ensuring the consistency of the captured images in time and space. In the tracheal intubation scene, these cameras can record the information of the target region from multiple perspectives, providing high-quality input data for subsequent stereo matching and depth calculation. Multi-perspective acquisition not only captures the surface structure and depth difference of the target region, but also lays a solid foundation for three-dimensional reconstruction and target recognition, thereby significantly improving the accuracy and reliability of image processing.

[0023] Step S120, the plurality of images are combined two by two to obtain image combinations. Specifically, on the basis of multi-view image acquisition, the plurality of acquired images are combined two by two to generate all possible image combinations. Specifically, assuming that n images are acquired by the multi-view camera, through two-by-two combination, n(n-1) / 2 image pairing results can be formed, each combination contains a pair of images taken from different perspectives. This way ensures comprehensive coverage of all perspective pairs, providing sufficient data support for subsequent stereo matching and depth calculation. At the same time, two-by-two combination optimizes the computational complexity, avoiding the high computational cost of directly processing multiple images, simplifying the processing logic and improving the accuracy on the basis of retaining complete perspective information, laying a solid foundation for subsequent three-dimensional reconstruction and target recognition.

[0024] ​Step S200, the at least one image combination is processed by stereo matching and depth calculation, and at least one depth map is obtained. Specifically, in the process of tracheal intubation, in order to obtain accurate depth information, the method processes the image combination by using a stereo matching algorithm, and deduces the depth information by analyzing the parallax of the pixels in the two images. First, the image combination is processed by grayscale processing to reduce the calculation complexity, and then processed by using a semi-global stereo matching algorithm (SGBM). The SGBM algorithm combines the advantages of local matching and global optimization, and is divided into three steps of matching cost calculation, path cost aggregation and disparity selection: a matching cost matrix is generated by the gray value difference, the path cost optimization result is aggregated along multiple directions, and finally the disparity value with the lowest cost is selected to generate a disparity map. Combined with the camera parameters (such as the focal length of the camera and the baseline distance) of the multi-view image acquisition unit, the depth value is calculated by the formula to generate a depth map pixel by pixel. The stereo matching algorithm not only effectively extracts the depth information in space, but also provides an accurate data basis for subsequent three-dimensional reconstruction and target recognition.

[0025] In one possible implementation, the at least one image combination is processed by stereo matching and depth calculation to obtain at least one depth map, and step S200 further includes step S210 of processing the at least one image combination by stereo matching based on a stereo matching algorithm to obtain at least one disparity map. Specifically, based on the stereo matching algorithm, the at least one image combination is processed to generate a disparity map. First, the images in the image combination are processed by grayscale processing to reduce the data dimension and retain key features such as edges and textures. Then, the gray value difference of the pixel points in the two images is compared by matching cost calculation, and the matching cost of the adjacent pixel points is aggregated to reduce noise, and an adaptive weighting method is used to realize cost aggregation. On this basis, the disparity value is optimized by using smooth constraint and uniqueness constraint, etc. to ensure the continuity and accuracy of the disparity map. Finally, the disparity value with the smallest cost is selected from the aggregation result to generate a disparity map, and the gray value in the disparity map represents the disparity size of the pixel, which can directly reflect the depth distribution of the objects in the scene, and provides an important basis for subsequent depth map generation and three-dimensional reconstruction.

[0026] Step S220, a plurality of camera parameters of the multi-view image acquisition unit are obtained, and at least one depth map is calculated and obtained in combination with the at least one disparity map. Specifically, in order to calculate the depth map, the camera parameters of the multi-view image acquisition unit need to be obtained first, including the focal length, optical center position and pixel size of the camera, and the baseline distance (i.e. the horizontal distance between the two cameras) between the cameras. These parameters are determined by the camera calibration process, and constitute the geometric model of the camera, providing a basis for depth calculation. Then, the camera parameters are combined with the disparity map for calculation. Each pixel value in the disparity map represents the horizontal displacement (parallax) of the scene point in the projection of the two images. Through the depth formula (Where Z is the depth, f is the focal length, B is the baseline distance, and d is the disparity value), the depth value of each point in the scene is calculated pixel by pixel, and a depth map is generated. Each pixel value in the depth map reflects the actual distance of the corresponding point to the camera, providing three-dimensional information of the scene. This process is of great significance in medical scenarios (such as tracheal intubation), providing real-time depth information for doctors to assist precise operations.

[0027] In one possible implementation, based on the stereo matching algorithm, the at least one image combination is subjected to stereo matching processing to obtain at least one disparity map, and step S210 further includes step S211, which subjects the at least one image combination to grayscale processing to obtain at least one grayscale image combination. Specifically, when processing the image combinations captured by the multi-view camera, the image combinations are first subjected to grayscale processing to simplify the subsequent calculation complexity and extract key image information. Grayscale is the process of converting a color image into a single-channel grayscale image. By removing color information, it focuses on the brightness component of the image, highlighting structural and edge features. The grayscale value is usually calculated using the weighted average method, and the range of the grayscale image after grayscale processing is 0 to 255, representing the gray levels from black to white. This process not only reduces the data dimension and enhances the structural information of the image, but also improves the efficiency and accuracy of stereo matching and depth calculation, laying a foundation for efficient and accurate subsequent processing.

[0028] Step S212, using the SGBM algorithm, the at least one grayscale image combination is subjected to stereo matching processing to obtain at least one disparity map. Specifically, in multi-view camera image processing, the SGBM (Semi-Global_Block_Matching) algorithm is used to perform stereo matching processing on the grayscale image combination to generate a disparity map. First, SGBM divides the input grayscale image into small blocks, searches for matching blocks in the left and right views through a sliding window, and calculates the initial matching cost based on the pixel brightness difference. Subsequently, SGBM optimizes the matching cost through multi-direction path cost accumulation (such as horizontal, vertical, and diagonal directions), smooths the noise effect, and reduces the calculation complexity. Finally, the algorithm selects the disparity value with the smallest cost as the final disparity of each pixel, thereby forming a disparity map. To improve the quality of the disparity map, SGBM also combines post-processing steps such as disparity filling, edge smoothing, and mismatch correction, making the disparity map more accurate and detailed. This process lays the foundation for subsequent depth calculation and three-dimensional reconstruction.

[0029] In a possible implementation, the plurality of camera parameters of the multi-view image acquisition unit are acquired, and at least one depth map is obtained by combining the at least one disparity map, and step S220 further includes step S221 of acquiring a plurality of camera focal lengths of the multi-view image acquisition unit. Specifically, in the multi-view camera image processing process, acquiring the focal length of each camera is a key basic step of depth calculation. The camera focal length (f) is defined as the distance from the optical center of the lens to the imaging plane, and is an important parameter affecting the accuracy of spatial depth perception. By performing a calibration experiment (such as using a chessboard calibration method) on each camera in the multi-view image acquisition unit, the intrinsic matrix of each camera can be extracted, and the focal length parameter can be read therefrom. If the camera has completed factory calibration, the focal length data can be directly obtained from the device configuration file. After the focal lengths of all cameras are acquired, the focal lengths need to be standardized to unify the unit and coordinate system and ensure the consistency of the parameters; for the focal length parameters with errors, the compensation algorithm can be used for correction. Finally, these focal length data are stored and used in the subsequent depth calculation process, to provide accurate optical basic data for each pair of camera combinations, and to ensure the accuracy and reliability of stereo matching and depth map generation.

[0030] Step S222, combining the multi-view image acquisition units according to the at least one image combination, and acquiring a baseline of two image acquisition units in each image acquisition unit combination to obtain at least one baseline. Specifically, in the multi-view camera image processing method, in order to calculate the image depth, the multi-view image acquisition units need to be combined according to the image combination rule, and the baseline between two image acquisition units in each group is acquired. Specifically, first, the cameras are grouped in pairs according to the configuration of the multi-view camera (such as a binocular or multi-view system), and each camera combination corresponds to an image pair that needs to be stereo matched. Then, the extrinsic matrix between each pair of cameras is acquired through the calibration experiment of the camera, which includes the distance between the optical centers of the two cameras, that is, the baseline (Baseline, B). The size of the baseline reflects the relative position relationship of the cameras in space, and is an important geometric parameter for depth calculation. For a multi-view camera system, such as a three-view camera system, the baseline values of each pair of camera combinations (such as cameras A and B, A and C, B and C) need to be calculated in sequence. The baseline data will be used in the subsequent depth map calculation to provide basic support for depth accuracy and measurable range. At the same time, reasonable design of the baseline length is of great significance to improve the system performance.

[0031] Step S223, at least one depth map is obtained by combining the at least one disparity map according to the plurality of camera focal lengths and at least one baseline, wherein the depth is calculated according to the following formula: wherein Z is the depth, f is the camera focal length, B is the baseline of the two image acquisition units, and d is the disparity. Specifically, in the multi-view camera image processing method, the depth calculation derives the three-dimensional depth information in the scene by combining camera parameters and a disparity map. Specifically, the camera focal length f and the baseline distance B between cameras, which represent the optical characteristics and spatial geometric relationship of the camera respectively, are utilized. At the same time, based on the disparity map d generated by the stereo matching algorithm, the disparity value of each pixel represents the difference in pixel position of the same spatial point in two images. Subsequently, the depth value Z is calculated pixel by pixel, wherein Z is the distance between the target point and the optical center of the camera. This formula is based on the principle of triangulation, accurately maps two-dimensional pixel information to three-dimensional space, generates a depth map, and each pixel point of the depth map corresponds to the depth information of a specific position in the scene. The generated depth map provides reliable three-dimensional spatial data support for subsequent three-dimensional reconstruction and target recognition tasks. The depth value Z is calculated pixel by pixel, wherein Z is the distance between the target point and the optical center of the camera. This formula is based on the principle of triangulation, accurately maps two-dimensional pixel information to three-dimensional space, generates a depth map, and each pixel point of the depth map corresponds to the depth information of a specific position in the scene. The generated depth map provides reliable three-dimensional spatial data support for subsequent three-dimensional reconstruction and target recognition tasks.

[0032] Step S300, according to the at least one depth map and multiple images, grayscale processing is performed to obtain at least one disparity map, three-dimensional reconstruction is performed to obtain a three-dimensional reconstruction map group. Specifically, in the multi-view camera image processing method, three-dimensional reconstruction realizes the three-dimensional modeling of the scene by combining the depth map and multiple images, and generates a three-dimensional reconstruction map group containing spatial information. First, according to the depth information of each pixel point in the depth map, the pixel points in the multiple images are marked as point cloud data with specific spatial position information. Then, by identifying the depth of each image pixel, multiple independent three-dimensional reconstruction maps are generated, which respectively present the scene structure characteristics under different viewing angles. Subsequently, by fusing these three-dimensional reconstruction maps, the information of different viewing angles is integrated, redundant or overlapping data is removed, and a complete three-dimensional reconstruction map group is generated. Finally, the map group accurately reflects the three-dimensional structure of the scene, providing a high-precision spatial basis for subsequent target recognition and depth analysis.

[0033] In a possible implementation, according to the at least one depth map and the plurality of images, a grayscale processing is performed to obtain at least one disparity map, and a three-dimensional reconstruction is performed to obtain a three-dimensional reconstruction group, and step S300 further includes step S310, in which depths of each pixel point in the at least one depth map are used to perform depth labeling on the plurality of images, and a three-dimensional reconstruction is completed to obtain a plurality of three-dimensional reconstruction maps. Specifically, in the three-dimensional reconstruction process, the depth map is used as a core data source, depth information of each pixel point in the depth map is used to perform depth labeling on the plurality of images, and pixel points in a two-dimensional image are mapped to a three-dimensional space to generate point cloud data having a real space position. By processing each image frame by frame and combining corresponding depth information, a plurality of three-dimensional reconstruction maps are respectively generated, the reconstruction maps display geometric features of a scene from different perspectives, and the accuracy of the depth information and the spatial consistency of the images are ensured. Finally, the three-dimensional reconstruction maps lay a foundation for complete three-dimensional reconstruction of the scene, fully restore the three-dimensional structure of the scene, and provide high-precision data support for subsequent integration analysis and target recognition.

[0034] Step S320, the plurality of three-dimensional reconstruction maps are combined to obtain a three-dimensional reconstruction group. Specifically, in the three-dimensional reconstruction process, the three-dimensional reconstruction maps generated from multiple perspectives are combined to form a complete three-dimensional reconstruction group. First, point cloud data of each three-dimensional reconstruction map is standardized in format to ensure consistency in a spatial coordinate system; then, a geometric feature matching algorithm (such as a common view point registration technology) is used to determine relative position relationships and rotation transformation parameters between three-dimensional reconstruction maps of different perspectives, and the three-dimensional reconstruction maps are spatially aligned. Next, the three-dimensional reconstruction maps of different perspectives are seamlessly fused by using a point cloud splicing technology, errors in edge overlapping areas are eliminated, and smoothness of data connection is ensured. Finally, the three-dimensional reconstruction maps of all perspectives are integrated into a complete three-dimensional reconstruction group, a three-dimensional structure of a target scene is fully displayed, and high-precision spatial data support is provided for subsequent target recognition and depth analysis.

[0035] Step S400, in the three-dimensional reconstruction map set, target recognition is performed, and the obtained target depth is output as the image processing result. Specifically, in the three-dimensional reconstruction map set, the target is recognized by a YOLO (You_Only_Look_Once) algorithm, and the target depth value is output as the image processing result. YOLO is a target detection algorithm based on deep learning, which can simplify the target detection task into a single regression problem, and can complete the classification and positioning of the target through a single forward propagation. Specifically, the plurality of three-dimensional reconstruction maps in the three-dimensional reconstruction map set are input into the YOLO target recognition channel one by one, the algorithm divides the image into a plurality of grids, and each grid predicts the class, boundary box position and confidence score of the target. After the recognition is completed, a plurality of basic depth values of the target region are extracted from the three-dimensional reconstruction map set, and these depth values correspond to the spatial position distribution of the target pixel points. In order to improve the accuracy of the depth values, the extracted basic depth values are subjected to mean value calculation, thereby obtaining stable and reliable target depth values. Finally, the target depth value is output as the result of image processing, realizing accurate recognition and positioning of the target, and improving the operation accuracy and efficiency in the tracheal intubation scene.

[0036] In a possible implementation, in the three-dimensional reconstruction map set, target recognition is performed, and the obtained target depth is output as the image processing result, and step S400 further includes step S410 of constructing a target recognition channel based on YOLO. Specifically, the process of constructing the target recognition channel based on the YOLO (You_Only_Look_Once) algorithm first utilizes the characteristics of YOLO to convert the target detection task into a regression problem of a single neural network, and outputs the boundary box position, size and class probability of the target through a single forward propagation. The channel takes the plurality of three-dimensional reconstruction maps in the three-dimensional reconstruction map set as input, extracts global features through a deep convolutional neural network (such as Darknet), and captures the spatial structure information of the target layer by layer. The input image is divided into fixed-size grids, and each grid predicts a plurality of boundary box parameters (such as center point coordinates, width, height and confidence score) and target class probability distribution. In order to improve the recognition effect, the channel defines a joint loss function including boundary box regression loss, confidence loss and class loss, and optimizes the network parameters through back propagation. For the tracheal intubation scene, the channel is trained and adapted on a specially labeled data set to ensure efficient prediction of the target position, size and class. Finally, the target recognition channel outputs the boundary box, classification result and confidence of the target, providing reliable support for subsequent depth analysis.

[0037] In step S420, the multiple three-dimensional reconstruction images in the three-dimensional reconstruction image group are input into the target recognition channel to identify the target. Specifically, the multiple three-dimensional reconstruction images in the three-dimensional reconstruction image group are input into the target recognition channel one by one, and the deep convolutional neural network constructed based on the YOLO (You_Only_Look_Once) algorithm is used to efficiently identify the target. The target recognition channel extracts high-level semantic features of the image, such as edges, textures, and shapes, through the feature extraction layer, and then divides the input image into grids in the detection layer. Each grid predicts the bounding box parameters (center point coordinates, width, height), confidence scores and category labels of multiple targets. In the tracheal intubation scenario, this channel can accurately identify relevant anatomical structures or catheter positions, and improve recognition accuracy through training with specific data sets. Finally, the target recognition channel outputs the bounding box, category label and confidence of the target, providing basic data support for subsequent depth calculation and positioning operations.

[0038] Step S430, extracting multiple basic depths of the target in the multiple three-dimensional reconstruction images. Specifically, the basic depth of the target is extracted in the multiple three-dimensional reconstruction images. First, the target recognition channel (such as a channel based on the YOLO algorithm) is used to locate the specific position of the target in each three-dimensional reconstruction image, including parameters such as the target's bounding box and center point. Subsequently, based on the bounding box area of ​​the target, the depth values ​​of all pixels in the area are extracted and recorded as the basic depth set of the target in the current three-dimensional reconstruction image. Since the target may exist in multiple three-dimensional reconstruction images at the same time, it is necessary to repeat this process in each image to extract the depth data sets of the target in different three-dimensional reconstruction images respectively. The extracted basic depth information depends on stereo matching and depth calculation processes to ensure the accuracy and reliability of the data, and provide high-quality support for subsequent depth mean calculation and depth characterization of the target.

[0039] Step S440, calculates the mean of the multiple basic depths and outputs it as the target depth as the image processing result. Specifically, after completing the extraction of the target basic depths in multiple three-dimensional reconstructed images, the mean of these depth data is calculated to finally obtain the comprehensive depth value of the target as the output result of the image processing. Specifically, first integrate the depth data of the target in each three-dimensional reconstructed image. These data are usually composed of pixel depth values ​​within the boundary area determined by the target recognition algorithm. After integration, perform statistical analysis on all depth data to eliminate possible outliers or noise interference to improve the accuracy of the results. Subsequently, perform arithmetic averaging on the depth data and directly calculate the mean of all basic depths. The calculation formula is: where Z avg represents the mean target depth, Z iFor the i-th basic depth value, N is the total number of depth values. Through this calculation method, the depth information of multiple perspectives can be integrated, and the errors caused by a single perspective can be eliminated, and finally accurate and stable target depth is output, providing reliable depth information support for subsequent tracheal intubation navigation or target positioning and other applications.

[0040] In the tracheal intubation process, the multi-view image acquisition unit collects multiple images and combines them to form an image pair. The image pair is subjected to stereo matching and depth calculation to generate a depth map. The depth map and the original image are combined for three-dimensional reconstruction to generate a three-dimensional model. Finally, the target is identified in the three-dimensional model and its depth is output as the image processing result, achieving the technical effect of improving the depth perception accuracy and real-time processing capability of the multi-view camera in tracheal intubation image processing.

[0041] In the foregoing, with reference to Figure 1 The intelligent multi-view camera tracheal intubation image processing method according to the embodiments of the present application is described in detail. Next, with reference to Figure 2 The intelligent multi-view camera tracheal intubation image processing device according to the embodiments of the present application will be described.

[0042] The intelligent multi-view camera tracheal intubation image processing device according to the embodiments of the present application solves the technical problems of inaccurate depth perception and insufficient real-time performance of existing multi-view cameras in tracheal intubation image processing, and achieves the technical effect of improving the depth perception accuracy and real-time processing capability of the multi-view camera in tracheal intubation image processing. The intelligent multi-view camera tracheal intubation image processing device includes a multi-view image acquisition unit, and further includes an image combination acquisition module 10, a depth map acquisition module 20, a three-dimensional reconstruction image group acquisition module 30, and a target depth output module 40.

[0043] The image combination acquisition module 10 is configured to acquire multiple images through the multi-view image acquisition unit during tracheal intubation, and combine the multiple images to obtain at least one image combination.

[0044] The depth map acquisition module 20 is configured to perform stereo matching processing and depth calculation on the at least one image combination to obtain at least one depth map.

[0045] The three-dimensional reconstruction image group acquisition module 30 is configured to perform grayscale processing on the at least one depth map and the multiple images to obtain at least one disparity map, and perform three-dimensional reconstruction to obtain a three-dimensional reconstruction image group.

[0046] The target depth output module 40 is configured to identify a target in the three-dimensional reconstruction image group and output a target depth as an image processing result.

[0047] In the following, the specific configuration of the image combination obtaining module 10 will be described in detail. As described above, in the tracheal intubation, a plurality of images are acquired by the multi-view image obtaining unit, and at least one image combination is obtained by combination, the image combination obtaining module 10 further comprises: an image acquisition unit, the image acquisition unit is used for acquiring a plurality of images by the multi-view image obtaining unit; an image traversal combination unit, the image traversal combination unit is used for traversing and combining the plurality of images two by two to obtain an image combination.

[0048] In the following, the specific configuration of the depth map obtaining module 20 will be described in detail. As described above, the at least one image combination is subjected to stereo matching processing and depth calculation to obtain at least one depth map, the depth map obtaining module 20 further comprises: a disparity map obtaining unit, the disparity map obtaining unit is used for performing stereo matching processing on the at least one image combination based on a stereo matching algorithm to obtain at least one disparity map; a depth map calculation unit, the depth map calculation unit is used for obtaining a plurality of camera parameters of the multi-view image obtaining unit, and combining the at least one disparity map to calculate and obtain at least one depth map.

[0049] In the following, the specific configuration of the depth map obtaining module 20 will be described in detail. As described above, the at least one image combination is subjected to stereo matching processing and depth calculation to obtain at least one depth map, the depth map obtaining module 20 further comprises: a disparity map obtaining unit, the disparity map obtaining unit is used for performing stereo matching processing on the at least one image combination based on a stereo matching algorithm to obtain at least one disparity map; a depth map calculation unit, the depth map calculation unit is used for obtaining a plurality of camera parameters of the multi-view image obtaining unit, and combining the at least one disparity map to calculate and obtain at least one depth map.

[0050] In the following, the specific configuration of the depth map obtaining module 20 will be described in detail. As described above, the at least one image combination is subjected to stereo matching processing and depth calculation to obtain at least one depth map, the depth map obtaining module 20 further comprises: a disparity map obtaining unit, the disparity map obtaining unit is used for performing stereo matching processing on the at least one image combination based on a stereo matching algorithm to obtain at least one disparity map; a depth map calculation unit, the depth map calculation unit is used for obtaining a plurality of camera parameters of the multi-view image obtaining unit, and combining the at least one disparity map to calculate and obtain at least one depth map. In the following, the specific configuration of the depth map obtaining module 20 will be described in detail. As described above, the at least one image combination is subjected to stereo matching processing and depth calculation to obtain at least one depth map, the depth map obtaining module 20 further comprises: a disparity map obtaining unit, the disparity map obtaining unit is used for performing stereo matching processing on the at least one image combination based on a stereo matching algorithm to obtain at least one disparity map; a depth map calculation unit, the depth map calculation unit is used for obtaining a plurality of camera parameters of the multi-view image obtaining unit, and combining the at least one disparity map to calculate and obtain at least one depth map.

[0051] In the following, the specific configuration of the three-dimensional reconstruction image group obtaining module 30 will be described in detail. As described above, according to the at least one depth image and the plurality of images, the gray processing is performed, the at least one disparity image is obtained, the three-dimensional reconstruction is performed, and the three-dimensional reconstruction image group is obtained. The three-dimensional reconstruction image group obtaining module 30 further comprises: a three-dimensional reconstruction image obtaining unit, configured to perform the depth identification on the plurality of images by using the depth of each pixel point in the at least one depth image, complete the three-dimensional reconstruction, and obtain a plurality of three-dimensional reconstruction images; and a three-dimensional reconstruction image combination unit, configured to combine the plurality of three-dimensional reconstruction images to obtain the three-dimensional reconstruction image group.

[0052] In the following, the specific configuration of the target depth output module 40 will be described in detail. As described above, in the three-dimensional reconstruction image group, the target recognition is performed, and the target depth is obtained as the image processing result. The target depth output module 40 further comprises: a target recognition channel construction unit, configured to construct a target recognition channel based on YOLO; a target recognition unit, configured to input a plurality of three-dimensional reconstruction images in the three-dimensional reconstruction image group into the target recognition channel to recognize and obtain a target; a basic depth extraction unit, configured to extract a plurality of basic depths of the target in the plurality of three-dimensional reconstruction images; and an image processing result output unit, configured to calculate the mean value of the plurality of basic depths, and output the mean value as the target depth as the image processing result.

[0053] The intelligent multi-view camera trachea cannula image processing device provided in the embodiments of the present application can execute the intelligent multi-view camera trachea cannula image processing method provided in any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0054] Figure 3 FIG. 1 is a structural schematic diagram of a system provided by an embodiment of the present application, showing a block diagram of an exemplary system suitable for implementing the embodiments of the present application. Figure 3 The system shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present application. The system is in the form of a general computing device, and its components can include but are not limited to an input device 301, a processor 302, a memory 303, and an output device 304. The processor 302 can be one or more; the memory 303 can include a computer readable medium and at least one program product, which has a set of (at least one) program modules configured to perform the functions of the embodiments of the present application.

[0055] The memory 303 shown in the embodiments of the present application can adopt any combination of one or more computer readable media; the computer readable storage medium can be, but is not limited to, an infrared ray, a semiconductor system, a device or a means, or any combination of the above, for storing software programs, computer executable programs and modules, such as the program instructions / modules corresponding to the intelligent multi-lens camera tracheal intubation image processing method in the embodiments of the present application. The processor 302 executes various functional applications and data processing of the computer device by running the software programs, instructions and modules stored in the memory 303, that is, implements the intelligent multi-lens camera tracheal intubation image processing method described above.

[0056] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or server, and the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be implemented; in addition, the specific names of the functional units are only for easy mutual differentiation, and do not limit the protection scope of the present application.

[0057] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A smart multi-lens camera tracheal intubation image processing method, characterized in that, The method is applied to an artificial intelligence multi-view camera tracheal intubation image processing device, the device comprising a multi-view image acquisition unit, the method comprising: In the tracheal intubation, a plurality of images are collected by the multi-view image acquisition unit, and at least one image combination is obtained by combination; At least one depth map is obtained by performing stereo matching processing and depth calculation on the at least one image combination; According to the at least one depth map and the plurality of images, at least one disparity map is obtained by performing grayscale processing, and a three-dimensional reconstruction map set is obtained by performing three-dimensional reconstruction. Target recognition is performed in the three-dimensional reconstruction map set, and target depth is output as an image processing result.

2. The artificial intelligence multi-view camera tracheal intubation image processing method of claim 1, wherein, The plurality of images are collected by the multi-view image acquisition unit, and at least one image combination is obtained by combination, comprising: The plurality of images are collected by the multi-view image acquisition unit; The plurality of images are combined by two-by-two traversal to obtain image combinations. 3.The artificial intelligence multi-view camera tracheal intubation image processing method of claim 1, wherein, The at least one depth map is obtained by performing stereo matching processing and depth calculation on the at least one image combination, comprising: At least one disparity map is obtained by performing stereo matching processing on the at least one image combination based on a stereo matching algorithm; A plurality of camera parameters of the multi-view image acquisition unit are obtained, and at least one depth map is calculated based on the plurality of camera parameters and the at least one disparity map.

4. The artificial intelligence multi-view camera tracheal intubation image processing method of claim 3, wherein, At least one disparity map is obtained by performing stereo matching processing on the at least one image combination based on a stereo matching algorithm, comprising: At least one grayscale map combination is obtained by performing grayscale processing on the at least one image combination; At least one disparity map is obtained by performing stereo matching processing on the at least one grayscale map combination using an SGBM algorithm. 5.The artificial intelligence multi-view camera tracheal intubation image processing method of claim 3, wherein, A plurality of camera parameters of the multi-view image acquisition unit are obtained, and at least one depth map is calculated based on the plurality of camera parameters and the at least one disparity map, comprising: A plurality of camera focal lengths of the multi-view image acquisition unit are obtained; At least one baseline is obtained by combining the multi-view image acquisition unit according to the at least one image combination and obtaining the baseline of two image acquisition units in each image acquisition unit combination. At least one depth map is calculated based on the plurality of camera focal lengths and at least one baseline and the at least one disparity map, wherein the depth is calculated according to the following formula: Wherein, Z is the depth, f is the camera focal length, B is the baseline of two image acquisition units, and d is the disparity.

6. The artificial intelligence multi-camera image processing method for tracheal intubation according to claim 1, wherein, At least one disparity map is obtained by performing grayscale processing on the at least one depth map and the plurality of images, and a three-dimensional reconstruction map set is obtained by performing three-dimensional reconstruction, comprising: A plurality of three-dimensional reconstruction maps are obtained by performing depth identification on the plurality of images using the depth of each pixel point in the at least one depth map to complete the three-dimensional reconstruction; The plurality of three-dimensional reconstruction maps are combined to obtain a three-dimensional reconstruction map set.

7. The artificial intelligence multi-camera image processing method for tracheal intubation according to claim 1, wherein, Target recognition is performed in the three-dimensional reconstruction map set, and target depth is output as an image processing result, comprising: A target recognition channel is constructed based on YOLO; The plurality of three-dimensional reconstruction maps in the three-dimensional reconstruction map set are input into the target recognition channel to identify the target; A plurality of basic depths of the target are extracted in the plurality of three-dimensional reconstruction maps; The mean value of the plurality of base depths is calculated, and the output is a target depth as an image processing result.

8. An intelligent multi-lens camera tracheal intubation image processing device, characterized in that, The device is used to implement the intelligent multi-view camera tracheal intubation image processing method of any one of claims 1-7, and the device comprises a multi-view image acquisition unit, and the device further comprises: An image combination acquisition module, which is used to acquire a plurality of images in tracheal intubation by the multi-view image acquisition unit and obtain at least one image combination by combination; A depth map acquisition module, which is used to perform stereo matching processing and depth calculation on the at least one image combination to obtain at least one depth map; A three-dimensional reconstruction map set acquisition module, which is used to perform grayscale processing on the at least one depth map and the plurality of images to obtain at least one parallax map, perform three-dimensional reconstruction, and obtain a three-dimensional reconstruction map set; A target depth output module, which is used to perform target recognition in the three-dimensional reconstruction map set and output a target depth obtained as an image processing result. 9.The artificial intelligence multi-view camera tracheal intubation image processing apparatus of claim 8, wherein, The image combination acquisition module comprises: An image acquisition unit, which is used to acquire a plurality of images by the multi-view image acquisition unit; An image traversal combination unit, which is used to perform two-by-two traversal combination on the plurality of images to obtain image combinations.

10. An intelligent multi-lens camera endotracheal tube image processing system, characterized in that, The system comprises: A memory, which is used to store executable instructions; A processor, which is used to execute the executable instructions stored in the memory to implement the intelligent multi-view camera tracheal intubation image processing method of any one of claims 1-7.

Citation Information

Cited By

  • Control method and system for binocular visual catheter with posture correction function

    CN121465492A

  • A control method and system for a binocular visual catheter with attitude correction

    CN121465492B