Electronic laryngoscope image recognition method and system based on artificial intelligence

CN122391157APending Publication Date: 2026-07-14HANGZHOU LION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU LION TECH CO LTD
Filing Date
2026-04-21
Publication Date
2026-07-14

Smart Images

  • Figure CN122391157A_ABST
    Figure CN122391157A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image recognition, and discloses an electronic laryngoscope image recognition method and system based on artificial intelligence. The method comprises the following steps: acquiring a continuous image sequence of a throat region, and recognizing a suspected lesion candidate region in the image through an identification model; acquiring a first image sequence, and judging whether each first candidate region in the first image exists in the sequence; if yes, calculating a motion vector consistency and rigidity degree parameter of the region based on the first image sequence, and judging whether the region is a lesion region according to the parameters; if no, acquiring a subsequent second image sequence, and judging whether the disappeared candidate region appears again; for the candidate region appearing again, calculating related parameters of the candidate region by comprehensively considering the first and second image sequences, and finally judging the candidate region by combining image features. The application can improve the accuracy and reliability of electronic laryngoscope image lesion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular to an image recognition method and system for electronic laryngoscopes based on artificial intelligence. Background Technology

[0002] With the development of artificial intelligence technology, deep learning-based image recognition models have been widely used in the assisted diagnosis of electronic laryngoscopy images, enabling rapid identification of suspected lesion areas such as tumors, leukoplakia, and keratosis. However, in actual clinical applications, the environment within the laryngeal cavity is complex, often containing interfering substances such as sputum, air bubbles, and food debris. These interfering substances may be highly similar to certain early lesions (such as ulcerative tumors and leukoplakia) in terms of visual features such as color and texture, leading to a large number of false alarms from the artificial intelligence recognition model, incorrectly labeling interfering areas as suspected lesions.

[0003] Existing identification methods typically rely solely on the static features of a single image, lacking analysis of the target's dynamic behavior over continuous time. For example, genuine lesions are usually fixed in location, move in coordination with normal laryngeal tissue, and have a relatively stable morphology; while interfering substances such as sputum and air bubbles may flow freely, disappear, or undergo drastic changes in shape. Therefore, relying solely on static image recognition is insufficient to eliminate these dynamic interferences, affecting the accuracy of lesion identification.

[0004] Therefore, the present invention provides an electronic laryngoscope image recognition method and system based on artificial intelligence. Summary of the Invention

[0005] This application provides an artificial intelligence-based electronic laryngoscope image recognition method and system to improve the accuracy of lesion recognition in laryngoscope images.

[0006] In a first aspect, this application provides an artificial intelligence-based electronic laryngoscope image recognition method, the method comprising: Step S1: Obtain a continuous image sequence of the laryngeal region, and identify suspected lesion candidate regions in the continuous image sequence based on a pre-trained recognition model; Step S2: Obtain the first image sequence in the continuous image sequence, obtain the first suspected lesion candidate region in the first image of the first image sequence, and determine whether each first candidate region in the first suspected lesion candidate region always exists in the first image sequence. Step S3: If yes, calculate the relevant parameters of the corresponding first candidate region based on the first image sequence. The relevant parameters include motion vector consistency and rigidity parameters. Determine whether the first candidate region is a lesion region based on the relevant parameters. If the first candidate region that has disappeared has appeared in the second image sequence, then proceed to step S5 if the first candidate region has disappeared in the second image sequence. If the first candidate region has disappeared in the second image sequence, then proceed to step S5 if the first candidate region has disappeared in the second image sequence. If the first candidate region has disappeared in the second image sequence, then proceed to step S5 if the first candidate region has disappeared in the second image sequence. Step S5: Calculate the relevant parameters of the corresponding first candidate region based on the first image sequence and the second image sequence, and determine whether the corresponding first candidate region is a real lesion region based on the relevant parameters and the corresponding image features.

[0007] In conjunction with the first aspect, in the first implementation of the first aspect of this application, calculating the motion vector consistency of the corresponding first candidate region based on the first image sequence includes: For each image in the first image sequence, select the closed contour of the corresponding first candidate region, calculate the average coordinate of all pixels within the closed contour in the current image, and obtain the first centroid coordinates of each first candidate region in each image; For each image in the first image sequence, identify the key regions that have always existed, obtain the corresponding second centroid coordinates of the key regions, and calculate the consistency of motion vectors based on the first centroid coordinates and the second centroid coordinates.

[0008] In conjunction with the first aspect, in the second implementation of the first aspect of this application, calculating the consistency of motion vectors based on the first centroid coordinates and the second centroid coordinates includes: A first coordinate system is preset. All first centroid coordinates in the first image sequence are connected into a first curve in the first coordinate system as the motion trajectory of the first candidate region. All second centroid coordinates are connected into a second curve in the first coordinate system as the motion trajectory of the key region. The correlation coefficient between the first curve and the second curve is calculated, and the correlation coefficient is used as the consistency of the motion vector.

[0009] In conjunction with the first aspect, in the third implementation of the first aspect of this application, the rigidity parameter of the corresponding first candidate region is calculated based on the first image sequence, including: Identify the anatomical feature regions of the vocal cords, determine the first and second images corresponding to the maximum opening and maximum closing phases in the cycle based on the spacing change curves of the feature regions, extract the geometric features of the first candidate region in the first image and the first candidate region in the second image, and calculate the rigidity parameter based on the geometric features.

[0010] In conjunction with the first aspect, in the fourth implementation of the first aspect of this application, the stiffness parameter is calculated based on geometric features, including: The geometric features of the first candidate region in the first image are taken as the first geometric features, and the geometric features of the first subsequent region in the second image are taken as the second geometric features. The rate of change of the first area of ​​the first geometric features and the second area of ​​the second geometric features is calculated as the first rigidity index. The rate of change of roundness is calculated based on the area and perimeter of the two geometric features, and the rate of change of roundness is taken as the second rigidity index. The rigidity parameter is calculated based on the first formula, which is: Where A is the rigidity parameter, a1 is the first rigidity index, and a2 is the second rigidity index.

[0011] In conjunction with the first aspect, in the fifth implementation of the first aspect of this application, determining whether the first candidate region is a lesion region based on relevant parameters includes: If the consistency of the motion vector in the relevant parameters is greater than the preset first threshold and the rigidity parameter is greater than the preset second threshold, the first candidate region is determined to be the lesion region; otherwise, the second candidate region is determined to be the interference region.

[0012] In conjunction with the first aspect, in the sixth implementation of the first aspect of this application, determining whether the corresponding first candidate region is a real lesion region based on relevant parameters and corresponding image features includes:

[0013] If the consistency of motion vectors in the relevant parameters is greater than a preset first threshold and the rigidity parameter is greater than a preset second threshold, then the corresponding image features are further obtained and further judgment is made in combination with the image features. If not, the first candidate region is determined to be an interference region.

[0014] In conjunction with the first aspect, in the seventh implementation of the first aspect of this application, further judgment is made based on image features, including: Obtain the target image in the image sequence where the corresponding first candidate region disappears. Obtain the preceding and following images of the target image. Based on the position of the first candidate region in the preceding and following images, predict the target position of the corresponding first candidate region in the target image. Obtain the brightness value of the target position in the target image. If the brightness value is greater than or equal to a preset brightness threshold, determine that the corresponding first candidate region is a lesion region; otherwise, determine that the first candidate region is an interference region.

[0015] Secondly, this application provides an artificial intelligence-based electronic laryngoscope image recognition system, the system comprising: The lesion identification module is used to acquire a continuous image sequence of the laryngeal region and identify suspected lesion candidate regions in the continuous image sequence based on a pre-trained identification model. The first judgment module is used to obtain a first image sequence in a continuous image sequence, obtain a first suspected lesion candidate region in the first image in the first image sequence, and determine whether each first candidate region in the first suspected lesion candidate region has always existed in the first image sequence. The second judgment module is used to calculate relevant parameters of the corresponding first candidate region based on the first image sequence when the condition is met. The relevant parameters include motion vector consistency and rigidity parameters. Based on the relevant parameters, the module determines whether the first candidate region is a lesion region. The third judgment module is used to obtain the second image sequence in the continuous image sequence if no, and to determine whether the first candidate region that has disappeared has appeared in the second image sequence. If yes, step S5 is executed; otherwise, the corresponding candidate region is determined to be a normal interference region. The fourth judgment module is used to calculate the relevant parameters of the corresponding first candidate region based on the first image sequence and the second image sequence, and to judge whether the corresponding first candidate region is a real lesion region based on the relevant parameters and the corresponding image features.

[0016] Compared with the prior art, the beneficial effects of the present invention are at least as follows: The technical solution provided in this application first obtains suspected lesion regions in a continuous image sequence based on a recognition model. Then, it determines whether each suspected lesion region persists in the sequence, initially screening out fixed candidate regions. For persistent candidate regions, two dynamic indicators, motion vector consistency and rigidity parameters, are introduced to compare the motion trajectory of the candidate region with that of surrounding normal tissue and its morphological changes during the vocal cord opening and closing cycle, accurately distinguishing real lesion regions from unstable morphological interferences such as sputum. For suspected lesion regions that have disappeared in the image sequence, a secondary judgment is made by tracking whether they reappear in subsequent image sequences and combining image brightness features, effectively identifying real lesion regions missed by the recognition model due to lens reflection, further improving the accuracy of lesion recognition. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of one embodiment of the electronic laryngoscope image recognition method based on artificial intelligence in this application. Figure 2This is a schematic diagram of one embodiment of the artificial intelligence-based electronic laryngoscope image recognition system in this application. Detailed Implementation

[0019] This application provides an artificial intelligence-based electronic laryngoscope image recognition method and system. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0020] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the artificial intelligence-based electronic laryngoscope image recognition method in this application includes: Step S1: Obtain a continuous image sequence of the laryngeal region, and identify suspected lesion candidate regions in the continuous image sequence based on a pre-trained recognition model; Specifically, with the development of artificial intelligence technology, AI recognition has been applied to lesion identification in electronic laryngoscope images. AI recognition models can quickly identify suspected lesion areas. Therefore, a continuous image sequence of the laryngeal region is acquired, and a pre-trained AI recognition model is used to identify the continuous image sequence and identify the suspected lesion candidate areas in each image of the image sequence. However, since sputum, bubbles, or food residue in the laryngeal cavity may be highly similar in color or texture to ulcerative tumors, leukoplakia, or keratosis, the lesion areas identified by AI may have errors. In order to further determine whether the suspected lesion candidate area is a real lesion area or an interference area (such as sputum, bubbles, etc.), subsequent steps are performed.

[0021] Step S2: Obtain the first image sequence in the continuous image sequence, obtain the first suspected lesion candidate region in the first image of the first image sequence, and determine whether each first candidate region in the first suspected lesion candidate region always exists in the first image sequence. Specifically, for true lesion areas, their location is usually fixed and does not disappear. However, for non-lesion areas, such as bubbles and debris, these interfering substances may appear in earlier image sequences but disappear in later image sequences. Therefore, in order to determine whether a suspected lesion area is a true lesion area, the first image sequence in the continuous image sequence is first obtained. Assuming there are 100 images in the image sequence, the first 50 images can be obtained as the first image sequence. Then, all suspected lesion candidate areas appearing in the first image of the first image sequence are taken as the first suspected lesion candidate areas. Each candidate area in the first suspected lesion candidate area is taken as the first candidate area, and it is determined whether it always exists in the first image sequence.

[0022] Step S3: If yes, calculate the relevant parameters of the corresponding first candidate region based on the first image sequence. The relevant parameters include motion vector consistency and rigidity parameters. Determine whether the first candidate region is a lesion region based on the relevant parameters. Specifically, for suspected lesion candidate regions that consistently exist in the first image sequence, it is determined that they are unlikely to disappear. Therefore, based on the first image sequence, relevant parameters of the corresponding first candidate region are calculated. These parameters include motion vector consistency and rigidity parameters. Motion vector consistency refers to the consistency between the motion vector of the first candidate region and the motion vector of normal tissue within a preset range around it. Since the motion of normal tissue is continuous, the motion of lesion tissue and normal tissue is roughly the same. However, the motion of interfering substances (such as sputum) is unstable and may change significantly with the subject's breathing or swallowing. Therefore, if the motion vector consistency between the two is high, it indicates that the corresponding first candidate region is a lesion region; otherwise, it indicates that the first candidate region may not be a lesion region. The rigidity parameter refers to the ability of the suspected lesion candidate region to maintain its shape when subjected to external force. Generally, lesion regions have a greater ability to maintain their shape, while interfering substances (such as sputum) may change significantly with the subject's breathing or swallowing. Therefore, the rigidity parameter is also calculated to determine whether the suspected lesion candidate region is a real lesion region or an interfering substance.

[0023] If the first candidate region that has disappeared has appeared in the second image sequence, then proceed to step S5 if the first candidate region has disappeared in the second image sequence. If the first candidate region has disappeared in the second image sequence, then proceed to step S5 if the first candidate region has disappeared in the second image sequence. If the first candidate region has disappeared in the second image sequence, then proceed to step S5 if the first candidate region has disappeared in the second image sequence. Specifically, for suspected lesion regions that are not consistently present in the first image sequence, they may be either interference or lesion areas. For interference such as bubbles, once they disappear, they are unlikely to reappear. However, there is another situation where lesion areas in the image sequence may be identified as reflective areas by the artificial intelligence recognition model in some image sequences due to reflection, leading to misjudgment. Since reflective areas may appear and disappear in consecutive image sequences as the lens moves, they may not be consistently present in the first image sequence. Therefore, in order to further accurately judge this situation, a second image sequence is obtained from the consecutive image sequence. The second image sequence is obtained after the first image sequence. Therefore, it is determined whether the candidate region that has disappeared has appeared in the second image sequence. If so, the corresponding first candidate region is further judged in subsequent steps. If not, it indicates that the corresponding first candidate region is interference (such as bubbles or residue), and therefore, the corresponding candidate region is judged as a normal interference region.

[0024] Step S5: Calculate the relevant parameters of the corresponding first candidate region based on the first image sequence and the second image sequence, and determine whether the corresponding first candidate region is a real lesion region based on the relevant parameters and the corresponding image features.

[0025] Specifically, to determine whether a previously disappeared suspected lesion area is a true lesion area, relevant parameters for the corresponding candidate area are calculated based on the first and second image sequences. The method for calculating these parameters is the same as that for calculating parameters based on the first image sequence. It's important to note that since the suspected lesion area has disappeared in some of the first and second image sequences, the images that disappeared are ignored when calculating the relevant parameters; only the images that appeared are used to calculate the parameters. Then, based on the relevant parameters and corresponding image features (such as whether the brightness value of the corresponding area at the location in the image where the lesion disappeared is a reflective area), the determination of whether the corresponding first candidate area is a true lesion area is made.

[0026] In one specific embodiment, calculating the motion vector consistency of the corresponding first candidate region based on the first image sequence specifically includes the following steps: For each image in the first image sequence, select the closed contour of the corresponding first candidate region, calculate the average coordinate of all pixels within the closed contour in the current image, and obtain the first centroid coordinates of each first candidate region in each image; For each image in the first image sequence, identify the key regions that have always existed, obtain the corresponding second centroid coordinates of the key regions, and calculate the consistency of motion vectors based on the first centroid coordinates and the second centroid coordinates.

[0027] Specifically, in order to calculate the consistency of motion vectors between the first candidate region and the normal tissue region, firstly, for each image in the first image sequence, the closed contour corresponding to the first candidate region is selected, and then the average coordinate of all pixels within the closed contour in the current image is calculated. The average coordinate is used as the first centroid coordinate of each first candidate region in each image. By calculating the centroid coordinate of the corresponding first candidate region for each image, the position of the first candidate region in each corresponding image can be obtained.

[0028] Then, for each image in the first image sequence, key regions adjacent to the first candidate region are identified. Key regions refer to normal tissues adjacent to the first candidate region that have special textures, such as key regions that can be detected in the vocal cord region in a continuous image sequence or other more recognizable key regions. The closed contours of the corresponding key regions are obtained, and the average coordinates of the corresponding closed contours in the current image are calculated as the second centroid coordinates of the key regions in each image.

[0029] The consistency of motion vectors can then be calculated based on the coordinates of the first and second centroids.

[0030] In one specific embodiment, the consistency of motion vectors is calculated based on the coordinates of the first centroid and the second centroid, specifically including the following steps: A first coordinate system is preset. All first centroid coordinates in the first image sequence are connected into a first curve in the first coordinate system as the motion trajectory of the first candidate region. All second centroid coordinates are connected into a second curve in the first coordinate system as the motion trajectory of the key region. The correlation coefficient between the first curve and the second curve is calculated, and the correlation coefficient is used as the consistency of the motion vector.

[0031] Specifically, the first centroid coordinate represents the center coordinate position of the first candidate region in its corresponding image, and the second centroid coordinate represents the center coordinate position of the key regions in their respective images. Since the key regions are normal tissue areas, their movement direction and amount change steadily and regularly as different images are taken over time. If the first candidate region is a lesion area, its movement direction and amount should be consistent with or similar to those of the key regions. If the first candidate region is an interference (such as sputum), its movement will have a significant sliding amount due to the slippage of sputum, resulting in a different movement trajectory from that of the key regions. Therefore, the first centroid coordinate is preset... A coordinate system is used to define the motion trajectory of the first candidate region by connecting all the first centroid coordinates in the first image sequence into a first curve, and the motion trajectory of the key region by connecting all the second centroid coordinates in the first coordinate system. The correlation coefficient between the first and second curves is calculated; this correlation coefficient can be the Pearson correlation coefficient. The correlation coefficient is used as the consistency of motion vectors. A large correlation coefficient indicates that the motion trajectories of the first candidate region and the key region are similar, suggesting that the first candidate region is more likely to be a lesion region. A small correlation coefficient indicates that the motion trajectories of the first candidate region and the key region are different, suggesting that the first candidate region is more likely to be an interference region. For more accurate judgment, a rigidity parameter is also considered to determine whether the first candidate region is a lesion region.

[0032] In one specific embodiment, calculating the rigidity parameter of the corresponding first candidate region based on the first image sequence includes the following steps: Identify the anatomical feature regions of the vocal cords, determine the first and second images corresponding to the maximum opening and maximum closing phases in the cycle based on the spacing change curves of the feature regions, extract the geometric features of the first candidate region in the first image and the first candidate region in the second image, and calculate the rigidity parameter based on the geometric features.

[0033] Specifically, the anatomical feature regions of the vocal cords (such as the bilateral lanceolates) in the first image sequence are located. The dynamic distance of the feature regions (the distance between the leftmost and rightmost sides of the bilateral lanceolates) is used as the spacing between the feature regions. Since the vocal cords exhibit regular opening and closing movements during respiration, the dynamic distance fluctuates sinusoidally over time. When the dynamic distance reaches a local maximum, it is marked as the maximum opening phase, and the corresponding image frame is marked as the first image. When the dynamic distance reaches a local minimum, it is marked as the maximum closing phase, and the corresponding image frame is marked as the second image. For each first candidate region, its geometric features in the first and second images are obtained. The geometric features include area and perimeter. The rigidity parameter is calculated based on the changes in the geometric features.

[0034] In one specific embodiment, the stiffness parameter is calculated based on geometric features, which specifically includes the following steps: The geometric features of the first candidate region in the first image are taken as the first geometric features, and the geometric features of the first candidate region in the second image are taken as the second geometric features. The rate of change of the first area of ​​the first geometric features and the second area of ​​the second geometric features is calculated as the first rigidity index. The rate of change of roundness is calculated based on the area and perimeter of the two geometric features, and the rate of change of roundness is taken as the second rigidity index. The rigidity parameter is calculated based on the first formula, which is: Where A is the rigidity parameter, a1 is the first rigidity index, and a2 is the second rigidity index.

[0035] Specifically, real lesions, such as vocal cord nodules or polyps, possess a certain degree of bioelasticity. Their area expands and contracts systematically with vocal cord stretching. Even under pressure and deformation, the overall connectivity and complexity of their structure remain relatively stable. In contrast, fluids like sputum undergo significant changes when compressed, leading to substantial alterations in area and shape. Therefore, the area change rate is used as the first rigid indicator, and the roundness change rate as the second rigid indicator. The roundness change rate is calculated based on the following formula: Where C1 is the first roundness of the corresponding first candidate region calculated based on the geometric features of the first image, and C2 is the second roundness of the corresponding first candidate region calculated based on the geometric features of the second image. The formula for calculating roundness is: Where C is roundness, S is area, and L is perimeter, the rigidity parameter is then calculated based on the first formula above. The larger the rate of change of area and the larger the rate of change of roundness, the smaller the calculated rigidity parameter. The rigidity parameter can then be used to determine whether the first candidate region is a lesion region.

[0036] In one specific embodiment, determining whether a first candidate region is a lesion region based on relevant parameters includes the following steps: If the consistency of the motion vector in the relevant parameters is greater than the preset first threshold and the rigidity parameter is greater than the preset second threshold, the first candidate region is determined to be the lesion region; otherwise, the second candidate region is determined to be the interference region.

[0037] Specifically, since the first candidate region has always existed in the first image sequence and has never disappeared, if its motion vector consistency is greater than the preset first threshold and its rigidity parameter is greater than the preset second threshold, it indicates that it meets the characteristics of fixed lesion tissue, and therefore the first candidate region is judged to be a lesion region; otherwise, it indicates that it does not meet the characteristics of fixed lesion tissue, and therefore it is judged to be an interference region.

[0038] In one specific embodiment, determining whether a first candidate region is a true lesion region based on relevant parameters and corresponding image features includes the following steps: If the consistency of motion vectors in the relevant parameters is greater than a preset first threshold and the rigidity parameter is greater than a preset second threshold, then the corresponding image features are further obtained and further judgment is made in combination with the image features. If not, the first candidate region is determined to be an interference region.

[0039] Specifically, since the first candidate region has disappeared in some images in the image sequence, there is a possibility that it is a reflective area and therefore the artificial intelligence model has not recognized it. Therefore, if the motion vector consistency parameter is greater than the preset first threshold and the rigidity parameter is greater than the preset second threshold, the corresponding image features are further obtained. The image features refer to the brightness features of the target image. Further judgment is made in combination with the image features. The specific method for further judgment will be explained in detail later. Otherwise, the corresponding first candidate region is judged not to meet the characteristics of fixed lesion tissue and is judged to be an interference region.

[0040] In one specific embodiment, further judgment is made based on image features, specifically including the following steps: Obtain the target image in the image sequence where the corresponding first candidate region disappears. Obtain the preceding and following images of the target image. Based on the position of the first candidate region in the preceding and following images, predict the target position of the corresponding first candidate region in the target image. Obtain the brightness value of the target position in the target image. If the brightness value is greater than or equal to a preset brightness threshold, determine that the corresponding first candidate region is a lesion region; otherwise, determine that the first candidate region is an interference region.

[0041] Specifically, since the first candidate region has disappeared in the image sequence, the target image corresponding to the disappearance of the first candidate region is acquired. This means the corresponding first candidate region in the target image was not identified. In this case, there is a possibility of misidentification of the corresponding region (e.g., a lesion of vitiligo, which is similar to a reflective area and therefore not identified as a suspected lesion). To avoid this, the preceding and subsequent images of the target image are acquired. The preceding image is the image taken before the target image, closest to the acquisition time of the target image, and contains the first candidate region. The subsequent image is the image taken after the target image, closest to the acquisition time of the target image, and also contains the first candidate region. Based on the positions of the first candidate region in the preceding and subsequent images, the target position of the corresponding first candidate region in the target image is predicted. For example, the first candidate region in the preceding image... The first position and the second position of the first candidate region in the subsequent image are used to predict the target position of the first candidate region in the target image. The brightness value of the target position in the target image is obtained. If the brightness value is greater than or equal to the preset brightness threshold, it means that the corresponding target position is very likely to be a reflective area in the target image, which leads to the lesion area not being identified. Since the first lesion area meets the characteristics of fixed lesion tissue, the corresponding first candidate region is judged to be a lesion area. Otherwise, it means that the corresponding target position is not a reflective area in the target image, so the possibility of it not being recognized by the recognition model is very small. However, since the corresponding first candidate region has disappeared, the possible flowable interference area (such as sputum, bubbles or residue) may have flowed to other positions when the corresponding target image was captured. Therefore, the first candidate region is judged to be an interference area.

[0042] The above describes the AI-based electronic laryngoscope image recognition method in the embodiments of this application. The following describes the AI-based electronic laryngoscope image recognition system in the embodiments of this application. Please refer to [link / reference]. Figure 2 One embodiment of the artificial intelligence-based electronic laryngoscope image recognition system in this application includes: The lesion identification module is used to acquire a continuous image sequence of the laryngeal region and identify suspected lesion candidate regions in the continuous image sequence based on a pre-trained identification model. The first judgment module is used to obtain a first image sequence in a continuous image sequence, obtain a first suspected lesion candidate region in the first image in the first image sequence, and determine whether each first candidate region in the first suspected lesion candidate region has always existed in the first image sequence. The second judgment module is used to calculate relevant parameters of the corresponding first candidate region based on the first image sequence when the condition is met. The relevant parameters include motion vector consistency and rigidity parameters. Based on the relevant parameters, the module determines whether the first candidate region is a lesion region. The third judgment module is used to obtain the second image sequence in the continuous image sequence if no, and to determine whether the first candidate region that has disappeared has appeared in the second image sequence. If yes, step S5 is executed; otherwise, the corresponding candidate region is determined to be a normal interference region. The fourth judgment module is used to calculate the relevant parameters of the corresponding first candidate region based on the first image sequence and the second image sequence, and to judge whether the corresponding first candidate region is a real lesion region based on the relevant parameters and the corresponding image features.

[0043] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0044] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0045] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. An artificial intelligence-based electronic laryngoscope image recognition method, characterized in that, The method includes: Step S1: Obtain a continuous image sequence of the laryngeal region, and identify suspected lesion candidate regions in the continuous image sequence based on a pre-trained recognition model; Step S2: Obtain the first image sequence in the continuous image sequence, obtain the first suspected lesion candidate region in the first image of the first image sequence, and determine whether each first candidate region in the first suspected lesion candidate region has always existed in the first image sequence. Step S3: If yes, calculate the relevant parameters of the corresponding first candidate region based on the first image sequence. The relevant parameters include motion vector consistency and rigidity parameters. Determine whether the first candidate region is a lesion region based on the relevant parameters. If the first candidate region that has disappeared has appeared in the second image sequence, then proceed to step S5 if the first candidate region has disappeared in the second image sequence. If the first candidate region has disappeared in the second image sequence, then proceed to step S5 if the first candidate region has disappeared in the second image sequence. If the first candidate region has disappeared in the second image sequence, then proceed to step S5 if the first candidate region has disappeared in the second image sequence. Step S5: Calculate the relevant parameters of the corresponding first candidate region based on the first image sequence and the second image sequence, and determine whether the corresponding first candidate region is a real lesion region based on the relevant parameters and the corresponding image features.

2. The method according to claim 1, characterized in that, The motion vector consistency of the corresponding first candidate region is calculated based on the first image sequence, including: For each image in the first image sequence, select the closed contour of the corresponding first candidate region, calculate the average coordinate of all pixels within the closed contour in the current image, and obtain the first centroid coordinates of each first candidate region in each image; For each image in the first image sequence, identify the key regions that have always existed, obtain the corresponding second centroid coordinates of the key regions, and calculate the consistency of motion vectors based on the first centroid coordinates and the second centroid coordinates.

3. The method according to claim 2, characterized in that, The consistency of motion vectors is calculated based on the coordinates of the first and second centroids, including: A first coordinate system is preset. All first centroid coordinates in the first image sequence are connected into a first curve in the first coordinate system as the motion trajectory of the first candidate region. All second centroid coordinates are connected into a second curve in the first coordinate system as the motion trajectory of the key region. The correlation coefficient between the first curve and the second curve is calculated, and the correlation coefficient is used as the consistency of the motion vector.

4. The method according to claim 1, characterized in that, The rigidity parameter of the corresponding first candidate region is calculated based on the first image sequence, including: Identify the anatomical feature regions of the vocal cords, determine the first and second images corresponding to the maximum opening and maximum closing phases in the cycle based on the spacing change curves of the feature regions, extract the geometric features of the first candidate region in the first image and the first candidate region in the second image, and calculate the rigidity parameter based on the geometric features.

5. The method according to claim 4, characterized in that, Stiffness parameters are calculated based on geometric features, including: The geometric features of the first candidate region in the first image are taken as the first geometric features, and the geometric features of the first candidate region in the second image are taken as the second geometric features. The rate of change of the first area of ​​the first geometric features and the second area of ​​the second geometric features is calculated as the first rigidity index. The rate of change of roundness is calculated based on the area and perimeter of the two geometric features, and the rate of change of roundness is taken as the second rigidity index. The rigidity parameter is calculated based on the first formula, which is: Where A is the rigidity parameter, a1 is the first rigidity index, and a2 is the second rigidity index.

6. The method according to claim 1, characterized in that, Determining whether the first candidate region is a lesion region based on relevant parameters includes: If the consistency of the motion vector in the relevant parameters is greater than the preset first threshold and the rigidity parameter is greater than the preset second threshold, the first candidate region is determined to be the lesion region; otherwise, the second candidate region is determined to be the interference region.

7. The method according to claim 1, characterized in that, Based on relevant parameters and corresponding image features, it is determined whether the first candidate region is a true lesion region, including: If the consistency of motion vectors in the relevant parameters is greater than a preset first threshold and the rigidity parameter is greater than a preset second threshold, then the corresponding image features are further obtained and further judgment is made in combination with the image features. If not, the first candidate region is determined to be an interference region.

8. The method according to claim 1, characterized in that, Further judgment is made based on image features, including: Obtain the target image in the image sequence where the corresponding first candidate region disappears. Obtain the preceding and following images of the target image. Based on the position of the first candidate region in the preceding and following images, predict the target position of the corresponding first candidate region in the target image. Obtain the brightness value of the target position in the target image. If the brightness value is greater than or equal to a preset brightness threshold, determine that the corresponding first candidate region is a lesion region; otherwise, determine that the first candidate region is an interference region.

9. An artificial intelligence-based electronic laryngoscope image recognition system, used to implement the artificial intelligence-based electronic laryngoscope image recognition method as described in any one of claims 1-8, characterized in that, The system includes: The lesion identification module is used to acquire a continuous image sequence of the laryngeal region and identify suspected lesion candidate regions in the continuous image sequence based on a pre-trained identification model. The first judgment module is used to obtain a first image sequence in a continuous image sequence, obtain a first suspected lesion candidate region in the first image in the first image sequence, and determine whether each first candidate region in the first suspected lesion candidate region has always existed in the first image sequence. The second judgment module is used to calculate relevant parameters of the corresponding first candidate region based on the first image sequence when the condition is met. The relevant parameters include motion vector consistency and rigidity parameters. Based on the relevant parameters, the module determines whether the first candidate region is a lesion region. The third judgment module is used to obtain the second image sequence in the continuous image sequence if no, and to determine whether the first candidate region that has disappeared has appeared in the second image sequence. If yes, step S5 is executed; otherwise, the corresponding candidate region is determined to be a normal interference region. The fourth judgment module is used to calculate the relevant parameters of the corresponding first candidate region based on the first image sequence and the second image sequence, and to judge whether the corresponding first candidate region is a real lesion region based on the relevant parameters and the corresponding image features.