Tracheal intubation guidance methods and guided laryngoscopes

By collecting images through laryngoscope and using the tracheal structure recognition model, visual and auditory feedback can be provided to guide non-professionals to perform tracheal intubation, solving the accuracy and efficiency problems of tracheal intubation in emergency treatment and improving the feasibility and safety of the operation.

CN120022488BActive Publication Date: 2025-09-12SECOND AFFILIATED HOSPITAL ZHEJIANG UNIV COLLEGE OF MEDICINE
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510515535.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-12
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In emergency situations, tracheal intubation is difficult because non-professionals find it difficult to accurately determine whether the tracheal structure is suitable for intubation. Especially when facing patients with infectious diseases, inexperienced medical staff face greater challenges, and existing technologies cannot effectively provide accurate tracheal intubation guidance.

Method used

By obtaining the image to be guided captured by the laryngoscope, the pre-trained tracheal structure recognition model is input for recognition, the area and position of the glottis structure are used to output intubation guidance information, and the operator is instructed to adjust the position of the laryngoscope through visual and auditory feedback to ensure the accuracy of intubation.

Benefits of technology

It significantly improves the accuracy and efficiency of tracheal intubation, helping non-professionals to perform precise tracheal intubation in emergency situations and reducing the risk of harm to patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120022488B_ABST
    Figure CN120022488B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of endoscopes, and discloses a tracheal intubation guidance method and a laryngoscope. The tracheal intubation guidance method comprises: obtaining an image to be guided for tracheal intubation guidance; inputting the image to be guided into a pre-trained tracheal structure recognition model to perform tracheal intubation structure recognition and obtain a recognition result; if the recognition result only contains a glottis structure, outputting intubation guidance information for indicating a first glottis exposure state and enabling tracheal intubation based on the glottis area and glottis area position of a target area corresponding to the glottis structure, or outputting correction guidance information for indicating a second glottis exposure state and instructing an operator to move the laryngoscope in a correction direction; if the recognition result contains other areas whose structural type is not a glottis structure, outputting correction guidance information based on the structural type corresponding to the other areas and / or the area corresponding to the other areas; the accuracy and efficiency of the tracheal intubation process can be significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of endoscopes, in particular to a tracheal intubation guiding method and a guiding laryngoscope. Background Art

[0002] Endotracheal intubation is a medical procedure that involves inserting a specially designed endotracheal tube through the mouth or nose, transglottis and into the trachea or bronchus to establish a reliable path for artificial ventilation.

[0003] In the current emergency medical system, non-professional medical personnel are often the first to arrive at the scene. For these non-professionals, intubation in emergency situations is challenging. Accurately assessing the suitability of the tracheal structure for intubation and making prompt clinical decisions are crucial. This skill typically requires extensive clinical experience and is difficult to acquire through conventional mannequin training alone.

[0004] Furthermore, it is worth noting that although more than half of anesthesiologists have received specialized training in difficult airway management, in some cases, such as intubating infectious patients, medical staff who lack sufficient experience face greater challenges due to the requirement to wear personal protective equipment. Therefore, strengthening training for airway management personnel is urgent.

[0005] Therefore, there is an urgent need to propose a tracheal intubation guidance method to accurately provide tracheal intubation guidance for professionals and non-professionals. Summary of the Invention

[0006] In view of this, the present invention provides an endotracheal intubation guidance method and a guiding laryngoscope to solve the problem of being unable to accurately provide endotracheal intubation guidance for professionals and non-professionals.

[0007] In a first aspect, the present invention provides a method for guiding tracheal intubation, comprising: obtaining an image to be guided for tracheal intubation; wherein the image to be guided is acquired by a laryngoscope; inputting the image to be guided into a pre-trained tracheal structure recognition model for tracheal intubation structure recognition to obtain a recognition result; if the recognition result only includes the glottis structure, based on the glottis area and glottis area position of the target area corresponding to the glottis structure, output intubation guidance information for indicating a first glottis exposure state and allowing tracheal intubation, or correction guidance information for indicating a second glottis exposure state and instructing the operator to move the laryngoscope in a correction direction; if the recognition result includes other areas whose structure type is not a glottis structure, outputting the correction guidance information based on the structure type corresponding to the other area and / or the area corresponding to the other area.

[0008] As an exemplary embodiment, the glottis area and glottis area position of the target area corresponding to the glottis structure are outputted to indicate a first glottis exposure state, and intubation guidance information for performing tracheal intubation, or correction guidance information for indicating a second glottis exposure state and instructing the operator to move the laryngoscope in a correction direction, including: obtaining the glottis area and glottis area position of the target area; if the area ratio of the glottis area in the image to be guided is greater than or equal to a first preset ratio, and the glottis area position satisfies the preset position, outputting the intubation guidance information; if the glottis area ratio of the glottis area relative to the image area of ​​the image to be guided is less than the first preset ratio and greater than the second preset ratio, outputting first correction guidance information; the first correction guidance information is used to indicate the second glottis exposure state and instruct to advance and lift the laryngoscope; if the glottis area position does not satisfy the preset position, outputting second correction guidance information based on the position difference between the glottis area and the preset position; wherein, the second correction guidance information is used to instruct to move the laryngoscope left or right.

[0009] As an exemplary embodiment, the outputting of correction guidance information based on the structural type corresponding to the other region and / or the area corresponding to the other region includes: obtaining the other structural type corresponding to the other region; if the other structural type does not include the epiglottis, outputting third correction guidance information; the third correction guidance information is used to instruct the laryngoscope to advance forward; if the other structural type only includes the epiglottis, outputting the correction guidance information based on the area of ​​the epiglottis.

[0010] As an exemplary embodiment, the correction guidance information is output based on the area of ​​the epiglottis region, including: obtaining the epiglottis region area of ​​a first other region whose other structural type is the epiglottis; if the area ratio of the epiglottis region in the image to be guided is less than a third preset ratio, outputting fourth correction guidance information; the fourth correction guidance information is used to indicate the first other region and instruct to advance the laryngoscope forward; if the epiglottis area ratio is greater than or equal to the third preset ratio, outputting fifth correction guidance information; the fifth correction guidance information is used to indicate the first other region and instruct to retract the laryngoscope backward.

[0011] As an exemplary embodiment, the method for training the tracheal structure recognition model includes: obtaining airway images and / or airway videos as a preset data set; performing at least one of expansion processing, annotation processing and enhancement processing on the preset data set to obtain a training data set; inputting the training data set into a pre-built structure recognition model for model training, and during the model training process, learning the correspondence between the actual deep semantic features and actual shallow spatial features of the training data set and the annotation results until the model converges to obtain the tracheal structure recognition model.

[0012] As an exemplary embodiment, the preset data set is expanded, including: inputting the preset data set into a pre-built generative adversarial network, and expanding the preset data set based on the generative adversarial network to obtain an expanded data set.

[0013] As an exemplary embodiment, the enhancement processing of the preset data set includes: performing random color adjustment on the expanded data set to obtain a color enhanced data set; performing scale transformation on the expanded data set to obtain a scale enhanced data set; and using the color enhanced data set and the scale enhanced data set as the enhanced data set.

[0014] As an exemplary embodiment, the enhancement processing of the expanded data set further includes: randomly selecting a geometric area in each data in the enhanced data set as the geometric area to be replaced; obtaining the original position of the geometric area to be replaced in the original data; and selecting random data in the enhanced data set for replacement based on the original position and the geometric area to be replaced, so as to enhance the preset data set.

[0015] As an exemplary embodiment, the structure recognition model includes a cascaded input end, a backbone network, a feature fusion network and a prediction end; the input end is used to preprocess the image to be guided and input the preprocessed image into the backbone network; the backbone network is used to extract features from the preprocessed image to obtain the actual deep semantic features and actual shallow spatial features of the preprocessed image; the feature fusion network is used to fuse the actual deep semantic features and the actual shallow spatial features to obtain fused features; the prediction end is used to generate a multi-scale feature map corresponding to each of the structure types based on the fused features to obtain the recognition result.

[0016] In a second aspect, the present invention provides a guided laryngoscope, comprising: a video acquisition module and an image processing module; the video acquisition module and the image processing module are communicatively connected to each other; the video acquisition module is used to acquire a patient's airway image and transmit the airway image to the image processing module;

[0017] The image processing module includes: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the tracheal intubation guidance method of the above-mentioned first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0018] The present invention provides a tracheal intubation guidance method, comprising: obtaining an image to be guided for tracheal intubation guidance; wherein the image to be guided is acquired by a laryngoscope; inputting the image to be guided into a pre-trained tracheal structure recognition model to perform tracheal intubation structure recognition to obtain a recognition result; if the recognition result only contains a glottis structure, outputting intubation guidance information for indicating a first glottis exposure state and enabling tracheal intubation or correction guidance information for indicating a second glottis exposure state and instructing an operator to move the laryngoscope in a correction direction based on the glottis area and glottis area position of a target area corresponding to the glottis structure; if the recognition result contains other areas whose structural types are not glottis structures, outputting the correction guidance information based on the structural types corresponding to the other areas and / or the area corresponding to the other areas; performing tracheal intubation structure recognition on the image to be guided to obtain a recognition result, and further determining the current position of the laryngoscope based on the structural type, area, and position of the area contained in the recognition result to guide the operator, thereby significantly improving the accuracy and efficiency of the tracheal intubation process. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1 is a flow chart of a tracheal intubation guidance method according to an embodiment of the present invention;

[0021] Figure 2 It is a schematic diagram of the network structure of the YOLOv5 model. DETAILED DESCRIPTION

[0022] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.

[0023] According to an embodiment of the present invention, an embodiment of a tracheal intubation guidance method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0024] In this embodiment, a tracheal intubation guidance method is provided, which is suitable for a guided laryngoscope (hereinafter referred to as a laryngoscope) Figure 1 FIG. 1 is a flow chart of a tracheal intubation guiding method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0025] Step S101 : obtaining an image to be guided for tracheal intubation guidance; wherein the image to be guided is acquired through laryngoscope.

[0026] When performing tracheal intubation, the patient's airway image is usually collected through the video acquisition module of the laryngoscope, and then the collected image is used to determine whether the current tracheal structure is suitable for intubation; based on this, in this embodiment, the image to be guided for tracheal intubation guidance is obtained by collecting the image through the laryngoscope; as a possible implementation method, the patient's airway image is directly collected through the laryngoscope as the image to be guided; as another possible implementation method, the patient's airway video stream is collected through the laryngoscope, and then the video stream is processed, such as segmentation, to obtain the image to be guided.

[0027] Step S102: input the image to be guided into a pre-trained tracheal structure recognition model to perform tracheal intubation structure recognition to obtain a recognition result.

[0028] According to clinical trials of existing laryngoscopes, when the video acquisition module acquires the patient's airway image, it can usually capture structures with obvious characteristics such as the epiglottis, glottis, tonsils, and uvula; in this embodiment, the tracheal structure recognition model is trained specifically for the above structures, so as to obtain the above one or more tracheal structures and their corresponding areas actually contained in the image to be guided by inputting the image to be guided obtained by the laryngoscope into the tracheal structure recognition model.

[0029] In one embodiment, the tracheal structure recognition model can be implemented using a machine learning model or a deep learning model. The tracheal structure recognition model can be trained using airway images collected by a laryngoscope and predefined labels associated with the epiglottis, glottis, tonsils, uvula, etc.

[0030] Step S103: If the recognition result only includes the glottis structure, based on the glottis area and the glottis position of the target area corresponding to the glottis structure, output is used to indicate a first glottis exposure state, and intubation guidance information for performing tracheal intubation is output, or is used to indicate a second glottis exposure state, and correction guidance information instructs the operator to move the laryngoscope in a correction direction.

[0031] The structure in the airway used for tracheal intubation is usually the glottis, but the glottis moves regularly during normal human activities, and the operator needs to perform intubation in an appropriate state; and there are cases where the recognition result includes the glottis structure, but the glottis structure is not in the preset position for intubation; therefore, if tracheal intubation is performed when the trachea is in an inappropriate state or in an incorrect position, forced entry of the trachea or incorrect timing of entry often causes harm to the patient; therefore, if the recognition result only includes the glottis structure, it indicates that the laryngoscope is located at a peripheral position of the glottis structure at this time; in order to further determine whether the glottis can meet the requirements of tracheal intubation, the glottis area area and glottis area position of the target area corresponding to the glottis structure are output to indicate a first glottis exposure state, and intubation guidance information that tracheal intubation can be performed, or correction guidance information that indicates a second glottis exposure state, and instructs the operator to move the laryngoscope in a correction direction.

[0032] In one embodiment, when the glottis area and the glottis area position can both meet the requirements for tracheal intubation, the glottis area and the glottis area position of the target area corresponding to the glottis structure are output to indicate the first glottis exposure state, and intubation guidance information for tracheal intubation can be performed.

[0033] For example, the intubation guidance information may be implemented by displaying an identification of a region frame representing a target region of the glottis structure on the image to be guided.

[0034] For example, the intubation guidance information can be implemented in the form of auditory or visual feedback such as sound, light, etc., so as to provide auditory feedback or visual feedback to medical personnel wearing personal protective equipment for operation.

[0035] As a specific embodiment, the intubation guidance information is realized by displaying an identification of a region frame representing the target area of ​​the glottis structure on the image to be guided and an audio prompt; specifically, when the area and position of the glottis region can meet the requirements for tracheal intubation, the region frame used to represent the target area of ​​the glottis structure is displayed on the image to be guided, and the intubation audio prompt information "The glottis is well exposed and intubation can be performed" is output.

[0036] In another embodiment, when either or both of the glottis area and the glottis area position cannot meet the requirements for tracheal intubation, correction guidance information is output based on the glottis area and the glottis area position of the target area to indicate the second glottis exposure state and to instruct the operator to move the laryngoscope in a correction direction.

[0037] Exemplarily, the correction guidance information may be realized by displaying a region frame representing the target region of the glottis structure and an arrow indicating the moving direction on the image to be guided.

[0038] Exemplarily, the correction guidance information may be implemented in the form of auditory or visual feedback such as sound or light, so as to provide auditory feedback or visual feedback to medical personnel who are wearing personal protective equipment and performing operations.

[0039] As a specific embodiment, the correction guidance information is realized through the area frame of the target area, the arrow indicating the moving direction and the audio prompt; specifically, when either or both of the glottis area and the glottis area position cannot meet the requirements of tracheal intubation, the area frame of the target area for characterizing the glottis structure and the arrow indicating the moving direction are displayed on the image to be guided, and the correction audio prompt information of "The glottis exposure is poor, please move the laryngoscope in the first direction" is output; wherein the first direction is determined based on the glottis area and the glottis area position.

[0040] Step S104: If the recognition result includes other regions whose structural types are non-glottal structures, the correction guidance information is output based on the structural types corresponding to the other regions and / or the area corresponding to the other regions.

[0041] If the recognition result includes other areas whose structural types are non-glottal structures, it indicates that the laryngoscope is not near the glottis area at this time, and the operator needs to be guided so that the movement direction of the laryngoscope tends to the glottis area; since the relative positions of other structures in the airway such as the epiglottis, tonsils, uvula, etc. to the glottis structure are fixed, therefore, in this embodiment, the correction guidance information is output based on the structural type corresponding to the other area and / or the area corresponding to the other area, so as to instruct the operator to move the laryngoscope according to the correction direction through the correction guidance information.

[0042] The tracheal intubation guidance method of this embodiment obtains an image to be guided for tracheal intubation guidance, wherein the image to be guided is acquired by laryngoscope; the image to be guided is input into a pre-trained tracheal structure recognition model to perform tracheal intubation structure recognition and obtain a recognition result; if the recognition result only includes the glottis structure, intubation guidance information indicating a first glottis exposure state and that tracheal intubation can be performed is output based on the glottis area and glottis area position of the target area corresponding to the glottis structure, or correction guidance information indicating a second glottis exposure state and instructing the operator to move the laryngoscope in a correction direction; if the recognition result includes other areas whose structural type is not the glottis structure, the correction guidance information is output based on the structural type and / or area corresponding to the other areas. The above method performs tracheal intubation structure recognition on the image to be guided to obtain a recognition result, and further determines the current position of the laryngoscope based on the structure type, area, and position included in the recognition result to guide the operator, significantly improving the accuracy and efficiency of the tracheal intubation process.

[0043] As an exemplary embodiment, based on the glottis area and glottis area position of the target area corresponding to the glottis structure, intubation guidance information for indicating a first glottis exposure state and capable of performing tracheal intubation or correction guidance information for indicating a second glottis exposure state and instructing the operator to move the laryngoscope in a correction direction is output, including: obtaining the glottis area and glottis area position of the target area; if the glottis area ratio of the glottis area relative to the image area of ​​the image to be guided is greater than or equal to a first preset ratio, and the glottis area position meets the preset position, outputting the intubation guidance information; if the glottis area ratio of the glottis area relative to the image area of ​​the image to be guided is less than the first preset ratio and greater than the second preset ratio, outputting the first correction guidance information; the first correction guidance information is used to indicate the second glottis exposure state and instruct to advance and lift the laryngoscope forward.

[0044] The regional area of ​​the glottis structure under different exposure states is different relative to the image area of ​​the image to be guided. Therefore, in this embodiment, the glottis exposure condition can be determined by the ratio of the glottis area to the image area of ​​the image to be guided.

[0045] Furthermore, when performing tracheal intubation, the glottis should be located at the relative center of the image to be guided. Therefore, in this embodiment, the guidance information outputted is matched by both the glottis region position and the glottis region area to guide the operator.

[0046] In one embodiment, if the glottis area ratio relative to the image area of ​​the image to be guided is greater than or equal to a first preset ratio, and the glottis area position satisfies a preset position, indicating that both the glottis exposure condition and the glottis position meet the requirements for tracheal intubation, intubation guidance information indicating the first glottis exposure condition and that tracheal intubation is possible is output. The intubation guidance information is output.

[0047] In one embodiment, if the glottis area ratio relative to the image area of ​​the image to be guided is less than a first preset ratio and greater than a second preset ratio, indicating that the exposure condition of the glottis cannot meet the requirements of tracheal intubation, a first correction guidance information is output at this time; the first correction guidance information is used to indicate the second glottis exposure state, and to instruct the laryngoscope to be advanced and lifted.

[0048] As a possible implementation, the first correction guidance information is implemented by displaying a region frame for representing the target region of the glottis structure on the image to be guided, displaying a flashing upward arrow, and outputting a first correction audio prompt message of "poor glottis exposure, please advance and lift the laryngoscope";

[0049] As a possible implementation, the first preset ratio may take any value in the range of [0.3, 0.36]. For example, the first preset ratio is 0.33.

[0050] As a possible implementation, the second preset ratio may take any value in the range of [0.01, 0.03]. For example, the first preset ratio takes 0.02.

[0051] In one embodiment, if the position of the glottis area does not meet the preset position, second correction guidance information is output based on the position difference between the glottis area and the preset position; wherein the second correction guidance information is used to instruct to move the laryngoscope left or right.

[0052] The position of the glottis region is obtained by ensuring that the geometric center of the glottis region is at a position on the X-axis determined relative to the horizontal direction of the image to be guided. As a possible implementation method, the length of the image to be guided on the X-axis is set to 1, and the leftmost side of the image to be guided is set as the origin of the X-axis. If the length obtained by the projection of the geometric center of the glottis region on the X-axis is not within the interval of [0.35, 0.65], it is determined that the position of the glottis region does not meet the preset position.

[0053] Among them, as a possible implementation method, the second correction guidance information is realized by displaying an area box for the target area for representing the glottis structure on the image to be guided, displaying a flashing arrow to the left, and outputting a second correction audio prompt information of "The glottis exposure is poor, please move the laryngoscope to the left"; or by displaying an area box for the target area for representing the glottis structure on the image to be guided, displaying a flashing arrow to the right, and outputting a second correction audio prompt information of "The glottis exposure is poor, please move the laryngoscope to the right".

[0054] If the length of the projection of the geometric center of the glottis area on the X-axis is in the interval of (0, 0.35), the second correction guidance information indicating the laryngoscope should be moved to the left is output.

[0055] If the length of the projection of the geometric center of the glottis area on the X-axis is within the interval of (0.65, 1), the second correction guide information instructing to move the laryngoscope to the right is output.

[0056] If the length of the projection of the geometric center of the glottis area on the X-axis is within the interval of [0.35, 0.65], it is confirmed that the position of the glottis area meets the preset position.

[0057] The relative positions of other structures in the airway, such as the epiglottis, tonsils, and uvula, to the glottis structure are fixed; specifically, the epiglottis is located in a position similar to a transportation hub in the entire airway; before the laryngoscope enters the throat, a relatively obvious image of at least one of the uvula and / or tonsils will be captured; after the laryngoscope enters the throat, a relatively obvious epiglottis structure will be captured; when the laryngoscope is a certain distance away from the epiglottis, the non-glottis structures in the recognition result include the epiglottis; when the laryngoscope is located at the root of the epiglottis, the glottis structure appears vaguely in the image to be guided; therefore, in the present invention, when the recognition result includes not only the glottis, the position of the laryngoscope can be determined by whether the recognition result includes the epiglottis.

[0058] Therefore, as an exemplary embodiment, the non-glottal structure includes the epiglottis, and the correction guidance information is output based on the structure type corresponding to the other region and / or the area corresponding to the other region, including: obtaining the other structure type corresponding to the other region; if the other structure type does not include the epiglottis, outputting third correction guidance information; the third correction guidance information is used to instruct the laryngoscope to advance forward; if the other structure type only includes the epiglottis, outputting the correction guidance information based on the area of ​​the epiglottis.

[0059] In one embodiment, if the other structural type does not include the epiglottis, indicating that the laryngoscope is at a position at the front end of the epiglottis that includes at least one of the tonsils or the uvula, a third correction guidance information is output to instruct the laryngoscope to advance forward, so as to instruct the operator to advance the laryngoscope forward through the third correction guidance information.

[0060] In another embodiment, if the other structural type does not include the epiglottis, and the laryngoscope is in a position where the tonsils and uvula are not identified and is in front of the tonsils and uvula, a third correction guidance information is output to instruct the operator to push the laryngoscope forward through the third correction guidance information.

[0061] Among them, as a possible implementation method, the third revised guidance information is implemented by displaying a flashing upward arrow on the image to be guided, and outputting a third revised audio prompt information of "Please push the laryngoscope forward".

[0062] Exemplarily, since the area of ​​the epiglottis region is different when the laryngoscope is in different positions including the epiglottis region, if the other structure type includes the epiglottis, the correction guidance information may be output based on the area of ​​the epiglottis region.

[0063] Specifically, as an exemplary embodiment, outputting the correction guidance information based on the area of ​​the epiglottis region includes:

[0064] Obtain the area of ​​the epiglottis region of the first other region whose other structural type is the epiglottis; if the area ratio of the epiglottis region in the image to be guided is less than the third preset ratio, output fourth correction guidance information; the fourth correction guidance information is used to indicate the first other region and instruct to advance the laryngoscope forward; if the area ratio of the epiglottis region is greater than or equal to the third preset ratio, output fifth correction guidance information; the fifth correction guidance information is used to indicate the first other region and instruct to retract the laryngoscope backward.

[0065] Exemplarily, if the epiglottis area ratio is less than the third preset ratio, it indicates that the laryngoscope is in a position that includes the epiglottis but is far away from the airway, and the operator should be instructed to advance the laryngoscope forward to approach the glottis; at this time, the fourth corrected guidance information for instructing the laryngoscope to advance forward is output to instruct the operator to advance the laryngoscope forward through the fourth corrected guidance information.

[0066] As a possible implementation, the fourth revised guidance information is implemented by displaying an area box for representing the first other area of ​​the epiglottis on the image to be guided, displaying a flashing upward arrow, and outputting a fourth revised audio prompt information of "Please push the laryngoscope forward".

[0067] Exemplarily, if the epiglottis area ratio is greater than or equal to the third preset ratio, it indicates that the laryngoscope is in a position that includes the epiglottis but is close to the airway, and the operator should be instructed to retract the laryngoscope backward to approach the glottis; at this time, the fifth correction guidance information for instructing the retraction of the laryngoscope is output; and the operator is instructed to retract the laryngoscope backward through the fifth correction guidance information.

[0068] As a possible implementation, the fifth revised guidance information is implemented by displaying an area box for representing the first other area of ​​the epiglottis on the image to be guided, displaying a flashing downward arrow, and outputting a fifth revised audio prompt information of "Please move the laryngoscope backwards".

[0069] As a possible implementation method, the third preset ratio may take any value in the range of [0.2, 0.4]; illustratively, the third preset ratio is 0.3.

[0070] As an exemplary embodiment, the method for training the tracheal structure recognition model includes: obtaining airway images and / or airway videos as a preset data set; expanding the preset data set to obtain an expanded data set; annotating the expanded data set to obtain an annotation result; enhancing the expanded data set to obtain a training data set; inputting the training data set into a pre-built structure recognition model for model training, and during the model training process, learning the correspondence between the actual deep semantic features and the actual shallow spatial features of the training data set and the annotation results until the model converges to obtain the tracheal structure recognition model.

[0071] Exemplarily, the preset data set may be composed of airway images and airway videos, either alone or together; wherein, the airway image can be obtained by collecting image data under anatomical structure and clinical state, the airway video can be obtained by collecting video stream data under anatomical structure and clinical state, and the image data and video stream data under clinical state can be obtained by laryngoscope collection.

[0072] Among them, as a possible implementation method, after obtaining the video stream of the collected anatomical structure and clinical status, the video stream is processed, such as segmentation.

[0073] Exemplarily, the preset data set includes the glottis and epiglottis cartilage regions in a normal state.

[0074] Exemplarily, the preset data set includes the glottis and epiglottis cartilage regions in a normal state and an abnormal state such as a tumor, inflammation, etc.

[0075] Exemplarily, the preset data set includes the glottis and epiglottis cartilage regions in a normal state and in an abnormal state such as a tumor or inflammation under different lighting conditions.

[0076] Exemplarily, the preset data set includes the tonsil, uvula, glottis, and epiglottis regions in normal states and abnormal states such as tumors and inflammation under different lighting conditions.

[0077] Tracheal intubation image data in emergency scenarios is usually difficult to obtain, resulting in a small overall sample size. To address this issue, it is necessary to expand the preset dataset consisting of airway images and airway videos. Therefore, after obtaining the preset dataset, in order to increase the number of samples in the preset dataset, the preset dataset is expanded.

[0078] Furthermore, in order to obtain the airway image and the associated predefined labels for the epiglottis, glottis, tonsils, uvula, etc., the expanded preset data set is annotated.

[0079] For example, LabelImg software can be used to annotate the expanded preset data set; LabelImg is an open source graphic image annotation tool used to create boundaries / rectangular boxes (suitable for annotating the position and size of objects to be annotated) and polygon annotations (suitable for annotating objects of irregular shapes); specifically, Labelimg software is used to select and annotate the tonsils, uvula, epiglottis cartilage and glottis structure to obtain the annotation results.

[0080] For example, a labeling result using LabelImg software is as follows:

[0081] - ImageA.jpg

[0082] - PictureA.txt

[0083] -- 0; 0.505515; 0.475694; 0.562500; 0.437500

[0084] -- 1; 0.515165; 0.531944; 0.107537; 0.133333

[0085] - Image B.jpg

[0086] - PictureB.txt

[0087] -- 2; 0.332721; 0.434722; 0.161765; 0.291667

[0088] -- 3; 0.688879; 0.140972; 0.138787; 0.270833

[0089] The label of each image is stored in a txt file with the same name, where the first number in each line represents the category to which the label belongs. In the present invention, 0 represents epiglottis, 1 represents glottis, 2 represents uvula, and 3 represents tonsil; the four numbers after each line represent the coordinates of the four vertices selected by the box; it should be understood that other representation methods can also be used to implement the labeling of the data set.

[0090] Furthermore, in order to improve the generalization ability of the model, prevent overfitting, and enhance the robustness of the model when the preset data set is input into the model for model training, the expanded data set is enhanced to obtain a training data set.

[0091] After obtaining the training data set, the training data set is input into a pre-built structure recognition model for model training. During the model training process, the correspondence between the actual deep semantic features and actual shallow spatial features of the training data set and the annotation results is learned until the model converges to obtain the tracheal structure recognition model.

[0092] As a possible implementation, the pre-built structure recognition model can be the YOLOv5 model. YOLOv5 is a widely used object detection algorithm that can achieve fast object recognition while maintaining high detection accuracy. This high efficiency makes YOLOv5 the algorithm of choice in many real-time applications.

[0093] Figure 2 This is a schematic diagram of the network structure of the YOLOv5 model, such as Figure 2 As shown in the figure, the YOLOv5 architecture is divided into four main parts: input end, backbone network (Backbone), feature fusion network (Neck) and prediction end; during the model training process, the training data set is input through the input end, the input image undergoes preprocessing steps such as normalization and scaling, enters the backbone network (Backbone) for feature extraction, and then passes through the feature fusion network (Neck) for feature fusion, and finally obtains the prediction result through multi-scale prediction at the prediction end.

[0094] Among them, the input end is used to perform preprocessing steps such as standardization and scaling on the input data to adapt to the input requirements of the model.

[0095] The backbone network uses CSP Darknet 53 as the feature extraction network. This network is an improved version of Darknet53, introducing the Cross Stage Partial (CSP) structure to enhance feature representation capabilities and making extensive use of modules such as CBL and C3. The C3 module contains a large number of residual structures, which helps prevent gradient explosion and vanishing problems caused by excessive network depth, thereby improving model training stability and performance.

[0096] Among them, the feature fusion network (Neck) is implemented by combining the Feature Pyramid Network (FPN) and the Path Aggregation Network (PAN);

[0097] This combination effectively integrates deep semantic information with shallow spatial information. FPN transfers deep semantic information from top to bottom, giving high-level feature maps stronger semantic information, while PAN transfers low-level spatial information from bottom to top, giving low-level feature maps stronger spatial positioning information. The FPN+PAN structure allows feature maps at different levels to be fused, preserving rich semantic information while maintaining high resolution. This helps improve detection of objects of varying scales.

[0098] Among them, the prediction end adopts a multi-scale prediction mechanism, which outputs feature maps of three different scales respectively, and can predict large, medium and small targets.

[0099] For example, YOLOv5 uses formula (1) as the bounding box loss function:

[0100] (1)

[0101] Where B is the area of ​​the minimum bounding rectangle of the predicted box and the ground-truth bounding box, x is the union of the two boxes, and IOU is the intersection-over-union ratio of the two bounding boxes. Using the bounding rectangle method can better represent the overlap problem of the two boxes.

[0102] For example, YOLOv5 uses a binary cross entropy function as the loss function for classification and confidence scoring.

[0103] Based on the network structure of the above-mentioned YOLOv5 model, in this embodiment, the training data set is input into the YOLOv5 model for model training. During the model training process, the parameters of the model are continuously adjusted, and the correspondence between the actual deep semantic features and the actual shallow spatial features of each data in the training data set and the annotation results is learned until the model converges to obtain the tracheal structure recognition model, so as to train a tracheal structure recognition model for tonsils, uvula, epiglottis cartilage and glottis structure using the YOLOv5 model as a framework.

[0104] Based on this, as an exemplary embodiment, the preset data set is expanded to obtain an expanded data set, including: inputting the preset data set into a pre-built generative adversarial network, and expanding the preset data set based on the generative adversarial network to obtain an expanded data set.

[0105] To address the limited amount of training data, this paper uses an efficient generative adversarial network (GAN) to synthesize new endotracheal intubation images to expand the dataset. This technique is particularly suitable for increasing the training data required for the recognition of key structures such as the glottis.

[0106] A GAN is a deep learning model consisting of two main components: a generator (G) and a discriminator (D). The generator's task is to produce realistic synthetic images that closely resemble the distribution of real samples; the discriminator's task is to distinguish between generated images and real images. Through the adversarial learning process between these two networks, the generator gradually learns to produce more realistic images, thereby generating a new dataset that approximates the distribution of real samples.

[0107] During GAN training, the goal of the generator network G is to generate realistic images to deceive the discriminator network D; while the goal of the discriminator network D is to distinguish the images generated by the generator network G from real images. In this way, G and D form a dynamic game process.

[0108] After training the GAN network D to better distinguish between images generated by the network G and real images, the preset data set is input into the pre-built generative adversarial network, and the preset data set is expanded based on the generative adversarial network to obtain an expanded data set.

[0109] In one embodiment, after the expanded data set is obtained, the expanded data set is labeled to obtain a labeling result.

[0110] As an exemplary embodiment, the enhancing the preset data set includes: performing random color adjustment on the expanded data set to obtain a color enhanced data set; performing scale transformation on the expanded data set to obtain a scale enhanced data set; and using the color enhanced data set and the scale enhanced data set as enhanced data sets.

[0111] Exemplarily, the scale transformation of the expanded dataset can be achieved through cropping, flipping, rotation, reflection, translation, etc.; specifically, in the cropping process, a part of the image is randomly cropped to simulate tracheal intubation images under different perspectives; in the flipping process, the image is flipped horizontally or vertically to increase the diversity of the data; in the rotation process, the image is randomly rotated by a certain angle to adapt to different shooting angles; in the reflection process, the image is mirrored to generate new samples; in the translation process, the target structure is randomly translated in the image to simulate the position change in actual operation.

[0112] Exemplarily, random color adjustment of the expanded data set can be achieved through color transformation and brightness adjustment; specifically, during the color transformation process, the brightness, contrast, saturation and hue of the image are randomly adjusted to simulate different lighting conditions; during the brightness adjustment process, the brightness of the image is randomly adjusted to cope with different lighting conditions.

[0113] In order to further improve the generalization ability and robustness of the model and make it better suitable for application scenarios, as an exemplary embodiment, the preset data set is expanded to obtain an expanded data set, which also includes: in the enhanced data set, randomly selecting a geometric area in each data as the geometric area to be replaced; obtaining the original position of the geometric area to be replaced in the original data; based on the original position and the geometric area to be replaced, selecting random data in the enhanced data set for replacement to enhance the preset data set.

[0114] Exemplarily, the above process can be implemented by CutMix technology; specifically, CutMix is ​​a method for generating new samples by mixing multiple training samples; during the mixing process, it randomly selects a rectangular area from one image and replaces it with the area at the corresponding position in another image.

[0115] In a second aspect, an embodiment of the present invention further provides a guided laryngoscope, comprising: a video acquisition module and an image processing module; the video acquisition module and the image processing module are communicatively connected to each other; the video acquisition module is used to acquire an airway image of a patient and transmit the airway image to the image processing module;

[0116] The image processing module includes: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the tracheal intubation guidance method described in any one of the above embodiments by executing the computer instructions.

[0117] In one embodiment, the image processing module includes a processor 10, a communication interface, a memory 30 and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus.

[0118] memory for storing computer programs;

[0119] The processor is configured to implement the tracheal intubation guidance method of any of the above embodiments when executing the computer program stored in the memory.

[0120] Optionally, in this embodiment, the communication bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc.

[0121] The communication interface is used for communication between the above-mentioned computer device and other devices.

[0122] The memory may include RAM, or may include non-volatile memory, such as at least one disk memory. Alternatively, the memory may also be at least one storage device located away from the aforementioned processor.

[0123] The above-mentioned processor can be a general-purpose processor, which can include but is not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processing), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0124] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and this embodiment will not be described in detail here.

[0125] A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which can include: a flash drive, ROM, RAM, a magnetic disk or an optical disk, etc.

[0126] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0127] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing one or more computer devices (such as personal computers, servers, or network devices) to execute all or part of the steps of the method in the above embodiments.

[0128] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, there may be other division methods, such as combining or integrating multiple units or components into another system, or ignoring or not implementing some features. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interface, indirect coupling or communication connection of units or modules, and may be electrical or other forms.

[0129] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the purpose of the solution provided in this embodiment.

[0130] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0131] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0132] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A guided laryngoscope, characterized in that: include: A video acquisition module and an image processing module; the video acquisition module and the image processing module are communicatively connected to each other; The video acquisition module is used to acquire the patient's airway image and transmit the airway image to the image processing module; The image processing module includes: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the following endotracheal intubation guidance method, including: Acquiring an image to be guided for tracheal intubation guidance; wherein the image to be guided is acquired through laryngoscope; Inputting the image to be guided into a pre-trained tracheal structure recognition model to perform tracheal intubation structure recognition to obtain a recognition result; If the recognition result only includes the glottis structure, outputting, based on the glottis area and the glottis area position of the target area corresponding to the glottis structure, intubation guidance information indicating a first glottis exposure state and enabling tracheal intubation, or correction guidance information indicating a second glottis exposure state and instructing the operator to move the laryngoscope in a correction direction; including: Acquire the glottis area and glottis position of the target area; If the area ratio of the glottis region in the image to be guided is greater than or equal to a first preset ratio, and the position of the glottis region satisfies a preset position, outputting the intubation guidance information; If the area ratio of the glottis region in the image to be guided is less than the first preset ratio and greater than the second preset ratio, outputting first correction guidance information; the first correction guidance information is used to indicate the exposure status of the second glottis and instruct to advance and lift the laryngoscope; If the position of the glottis area does not meet the preset position, outputting second correction guidance information based on the position difference between the glottis area and the preset position; wherein the second correction guidance information is used to instruct to move the laryngoscope leftward or rightward; The position of the glottis region is obtained by locating the geometric center of the glottis region at the position of the X-axis determined relative to the horizontal direction of the image to be guided; the length of the image to be guided on the X-axis is set to 1, and the leftmost side of the image to be guided is set as the origin of the X-axis. If the length of the projection of the geometric center of the glottis region on the X-axis is in the interval of (0, 0.35), the second correction guidance information indicating that the laryngoscope should be moved to the left is output; if the length of the projection of the geometric center of the glottis region on the X-axis is in the interval of (0.65, 1), the second correction guidance information indicating that the laryngoscope should be moved to the right is output; if the length of the projection of the geometric center of the glottis region on the X-axis is in the interval of [0.35, 0.65], it is confirmed that the position of the glottis region meets the preset position; If the recognition result includes other areas whose structural types are non-glottal structures, the correction guidance information is output based on the structural types corresponding to the other areas and / or the area corresponding to the other areas; wherein the other areas include the uvula, tonsils and / or epiglottis; the correction guidance information is output based on the structural types corresponding to the other areas and / or the area corresponding to the other areas, including: obtaining the other structural types corresponding to the other areas; if the other structural types do not include the epiglottis, outputting third correction guidance information; the third correction guidance information is used to instruct the laryngoscope to be pushed forward; if the other structural types only include the epiglottis, outputting the correction guidance information based on the area of ​​the epiglottis area.

2. The guided laryngoscope according to claim 1, wherein: Outputting the correction guidance information based on the area of ​​the epiglottis region includes: Obtaining an area of ​​a first other region of the epiglottis where the other structural type is the epiglottis; if a proportion of the area of ​​the epiglottis in the image to be guided is less than a third preset proportion, outputting fourth corrected guidance information; the fourth corrected guidance information is used to indicate the first other region and instruct to advance the laryngoscope; If the area ratio of the epiglottis region is greater than or equal to the third preset ratio, fifth correction guidance information is output; the fifth correction guidance information is used to indicate the first other region and instruct to retract the laryngoscope backward.

3. The guided laryngoscope according to claim 1, wherein: The method for training the tracheal structure recognition model includes: Acquire airway images and / or airway videos as a preset data set; Performing at least one of expansion processing, labeling processing, and enhancement processing on the preset data set to obtain a training data set; The training data set is input into a pre-built structure recognition model for model training. During the model training process, the correspondence between the actual deep semantic features and actual shallow spatial features of the training data set and the annotation results is learned until the model converges to obtain the tracheal structure recognition model.

4. The guided laryngoscope according to claim 3, wherein: The expanding the preset data set includes: The preset data set is input into a pre-built generative adversarial network, and the preset data set is expanded based on the generative adversarial network to obtain an expanded data set.

5. The guided laryngoscope according to claim 4, wherein: Performing enhancement processing on the expanded data set, including: Performing random color adjustment on the expanded data set to obtain a color enhanced data set; Performing a scale transformation on the expanded data set to obtain a scale-enhanced data set; The color enhancement dataset and the scale enhancement dataset are used as enhancement datasets.

6. The guided laryngoscope according to claim 5, characterized in that: Performing enhancement processing on the expanded data set further includes: In the enhanced data set, randomly selecting a geometric region in each data as the geometric region to be replaced; Obtaining the original position of the geometric area to be replaced in the original data; Random data is selected from the enhanced data set for replacement based on the original position and the geometric area to be replaced, so as to enhance the preset data set.

7. The guided laryngoscope according to claim 3, wherein: The structure recognition model includes a cascade input end, a backbone network, a feature fusion network and a prediction end; The input end is used to preprocess the image to be guided and input the preprocessed image into the backbone network; The backbone network is used to extract features from the preprocessed image to obtain actual deep semantic features and actual shallow spatial features of the preprocessed image; The feature fusion network is used to fuse the actual deep semantic features and the actual shallow spatial features to obtain fused features; The prediction end is used to generate a multi-scale feature map corresponding to each of the structural types based on the fusion features to obtain the recognition result.

Citation Information

Patent Citations

  • Laryngoscope guide system and cannula guide subsystem

    CN115400310A

  • Glottis identification method and device and computer readable storage medium

    CN115713718A

  • Image processing device, method for operating same, and endoscope system

    CN117637120A

  • Method, system and equipment for identifying and positioning glottis epiglottis structure of human body based on universal embedded model with edge computing capability

    CN117974982A

  • Intelligent laryngoscope and use method

    CN118356140A