Intelligent identification and labeling method, system and equipment for endoscope image

By normalizing the brightness and suppressing specular reflection in endoscopic image sequences, extracting anatomical features, and performing local directional texture consistency analysis, the problems of subjectivity and inconsistent labeling in endoscopic image lesion identification are solved, achieving highly accurate lesion identification and structured description.

CN121837331AInactive Publication Date: 2026-04-10BASEMED (SHANGHAI) MEDICAL EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BASEMED (SHANGHAI) MEDICAL EQUIP CO LTD
Filing Date
2026-03-10
Publication Date
2026-04-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Current endoscopic image lesion identification and annotation rely on manual observation, which is easily affected by operator experience, fatigue and imaging conditions, leading to missed detection, false detection and inconsistent annotation of lesions. Moreover, existing technologies do not make full use of the temporal continuity information of image sequences, making it difficult to generate structured descriptions.

Method used

By performing brightness normalization and specular reflection suppression on endoscopic image sequences, anatomical structural features are extracted, anatomical structural constraint maps are generated, local directional texture consistency analysis is performed, abnormal texture perturbation components are established, and multi-scale aggregation and continuity analysis are conducted to construct significant candidate maps of lesions and continuous correction results.

Benefits of technology

It improves the consistency of endoscopic image quality and the accuracy of annotation, reduces the subjectivity and false detection rate of lesion identification, and generates structured descriptive information containing lesion location, morphology and characteristics, meeting the needs of clinical diagnosis and treatment records and teaching research.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121837331A_ABST
    Figure CN121837331A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent identification labeling method, system and device for an endoscope image, and relates to the technical field of image identification. The method comprises the steps that an endoscope image sequence is read, brightness normalization and reflection suppression are carried out, and a standard image sequence is generated; extracting an anatomical structure feature set based on features such as a cavity edge contour, and generating an anatomical structure constraint graph; performing local direction texture consistency analysis in a limited range, and establishing an abnormal texture disturbance component; and carrying out multi-scale aggregation on the disturbance components to establish a lesion candidate map, executing adjacent frame space position continuity analysis, calculating offset and form variation to construct a continuous correction result, and finally executing image annotation. The technical problems that the image quality is unstable and the image labeling precision is low in the continuous collection process of the endoscope image are solved, and the technical effect that the image quality consistency, stability and labeling accuracy are improved by fusing anatomical prior constraints and multi-frame dynamic continuity analysis is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition technology, and more specifically to intelligent recognition and annotation methods, systems and devices for endoscopic images. Background Technology

[0002] Endoscopy, as an important tool for screening and diagnosing digestive tract diseases, has been widely used in the clinical detection of gastrointestinal tumors, inflammatory diseases, and early lesions. With the continuous improvement of endoscopic imaging resolution and acquisition speed, a large amount of continuous image or video data can be generated during a single examination, which places higher demands on doctors' efficiency in interpreting images and their diagnostic accuracy.

[0003] Currently, lesion identification and annotation in endoscopic images mainly rely on manual observation and experience. Doctors need to identify and record mucosal morphology, texture changes, and abnormal areas in real time during the dynamic examination process. This method is not only labor-intensive and highly subjective, but also easily affected by the operator's experience level, fatigue level, and changes in imaging conditions (such as uneven lighting, specular reflection, motion blur, etc.), leading to problems such as missed lesions, false positives, and inconsistent annotations. It has significant limitations, especially in the identification of early, small lesions and lesions with abnormal texture.

[0004] To improve automation, some existing technologies attempt to introduce image processing or deep learning methods for lesion detection in endoscopic images. However, these methods often focus on target recognition or classification in single-frame images, failing to fully utilize the temporal continuity information inherent in the image sequence during continuous endoscopic advancement. Furthermore, they do not adequately consider the complex anatomical background, the directionality of mucosal folds, and the periodicity of tissue texture within the endoscopic field of view, easily misidentifying normal structures as lesions or causing lesion location drift with changes in viewing angle, affecting the stability and clinical reference value of the annotation results. In addition, existing technologies primarily rely on simple bounding boxes or category labels for image annotation, making it difficult to automatically generate structured descriptive information including lesion location, shape, size, and endoscopic features, thus failing to meet the needs of standardized textual and graphic reports for clinical records, follow-up comparisons, and teaching research. Summary of the Invention

[0005] This application provides an intelligent recognition and annotation method, system, and device for endoscopic images, which solves the technical problems of unstable image quality and low image annotation accuracy during continuous acquisition of endoscopic images.

[0006] The first aspect of this application provides an intelligent recognition and annotation method for endoscopic images, the method comprising: The system reads the image sequence acquired during continuous endoscopy, performs brightness normalization and specular reflection suppression on each frame of the image sequence, and generates a standard endoscopic image sequence with consistent imaging conditions. Based on the cavity edge contour, mucosal fold direction, and tissue texture periodicity, a set of structural features is extracted to characterize the distribution of anatomical structures within the endoscopic field of view, and an anatomical constraint map is generated based on the set of structural features. Within the spatial range defined by the anatomical constraint map, local directional texture consistency analysis of the standard endoscopic image sequence is performed to establish abnormal texture perturbation components. The abnormal texture perturbation components are aggregated at multiple scales to establish significant lesion candidate maps, and the spatial position continuity analysis of adjacent frames of the significant lesion candidate maps is performed based on the standard endoscopic image sequence to calculate position offset and morphological change, and construct a continuous correction result. The continuous correction result is used to perform image annotation of the standard endoscopic image sequence.

[0007] A second aspect of this application provides an intelligent recognition and annotation system for endoscopic images, the system comprising: Image standardization module: Reads the image sequence acquired during continuous endoscopy, performs brightness normalization and specular reflection suppression processing on each frame of the image sequence, and generates a standard endoscopic image sequence with consistent imaging conditions; Feature extraction module: Extracts a set of structural features to characterize the distribution of anatomical structures within the endoscopic field of view based on cavity edge contours, mucosal fold direction, and tissue texture periodicity features, and generates an anatomical structure constraint map based on the set of structural features; Consistency analysis module: Performs local directional texture consistency analysis on the standard endoscopic image sequence within the spatial range defined by the anatomical structure constraint map, and establishes abnormal texture perturbation components; Continuity analysis module: Performs multi-scale aggregation on the abnormal texture perturbation components, establishes significant lesion candidate maps, and performs spatial position continuity analysis of adjacent frames of significant lesion candidate maps based on the standard endoscopic image sequence, calculates position offset and morphological change, and constructs continuous correction results; Image annotation module: Performs image annotation on the standard endoscopic image sequence using the continuous correction results.

[0008] A third aspect of this application provides an electronic device, comprising: a memory for storing executable instructions; and a processor for executing the executable instructions stored in the memory to implement the intelligent recognition and annotation method for endoscopic images provided in this application.

[0009] One or more technical solutions provided in this application have at least the following technical effects or advantages: This method involves reading image sequences acquired during continuous endoscopic advancement, performing brightness normalization and specular reflection suppression on each frame to generate a standard endoscopic image sequence with consistent imaging conditions. Subsequently, based on cavity edge contours, mucosal fold orientation, and tissue texture periodicity, a set of structural features characterizing the distribution of anatomical structures within the endoscopic field of view is extracted, and an anatomical constraint map is generated based on this set. Then, within the spatial range defined by the anatomical constraint map, local directional texture consistency analysis of the standard endoscopic image sequence is performed to establish abnormal texture perturbation components. Next, multi-scale aggregation of these abnormal texture perturbation components is performed to establish salient lesion candidate maps. Based on the standard endoscopic image sequence, spatial position continuity analysis of adjacent frames of these candidate maps is performed, calculating positional offsets and morphological changes to construct continuous correction results. Finally, image annotation of the standard endoscopic image sequence is performed using these continuous correction results. This method solves the technical problems of unstable image quality and low annotation accuracy during continuous endoscopic image acquisition, achieving the technical effect of improving image quality consistency, stability, and annotation accuracy by fusing anatomical prior constraints and multi-frame dynamic continuity analysis. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a schematic flowchart of an intelligent recognition and annotation method for endoscopic images provided in an embodiment of this application.

[0012] Figure 2 This is a schematic diagram of the intelligent recognition and annotation system for endoscopic images provided in an embodiment of this application.

[0013] Figure 3 This is a schematic diagram of the structure of an exemplary electronic device of this application.

[0014] Figure labeling: Image standardization module 11, feature extraction module 12, consistency analysis module 13, continuity analysis module 14, image annotation module 15, processor 21, memory 22, input device 23, output device 24. Detailed Implementation

[0015] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0016] Example 1, as Figure 1 As shown in the embodiments of this application, an intelligent recognition and annotation method for endoscopic images is provided, wherein the method includes: The image sequence acquired during the continuous advancement of the endoscope is read, and brightness normalization and specular reflection suppression processing are performed on each frame of the image sequence to generate a standard image sequence of the endoscope with consistent imaging conditions.

[0017] In one embodiment, after reading the video stream acquired during the continuous advancement of the endoscope, the video stream is first parsed at the frame level, and a sequence of original images arranged in chronological order is extracted according to a preset frame rate. For each frame of the original image, the effective field of view is first determined. That is, the black border areas around the image are removed by segmenting or fixing a circular field of view template using a brightness threshold, resulting in an effective imaging mask. Subsequent calculations will only be performed on the pixels within the mask. Subsequently, the effective imaging mask is converted to a brightness representation space, and the brightness values ​​of the pixels within the effective imaging area are statistically analyzed to calculate the mean, standard deviation, and upper and lower quantiles of brightness. Then, based on the preset target brightness distribution parameters, linear stretching or adaptive histogram equalization is performed on the brightness values ​​of the current frame to bring the mean brightness and dynamic range of the processed image to a unified target range, thereby eliminating inter-frame brightness fluctuations caused by changes in the depth of endoscope advancement and the distance to the light source. After brightness normalization, the system performs specular reflection suppression. Specifically, in the normalized image, pixels are jointly judged based on high brightness and low saturation thresholds. Pixels that simultaneously meet the criteria of brightness close to the upper limit and color saturation below the threshold are marked as candidate reflective pixels. Connectivity analysis is then performed on these candidate reflective pixels, and adjacent reflective regions are merged through morphological operations such as dilation and closing to generate a continuous reflective region mask. Next, for pixels within the reflective region mask, a neighborhood-guided reflection correction strategy is adopted. Using non-reflective pixels within a certain distance of the reflective region boundary as a reference, weighted interpolation reconstruction is performed on the brightness and color values ​​of pixels within the reflective region, gradually transitioning them to the texture and brightness distribution of the surrounding tissue region. Smooth transition processing is performed at the reflective region boundary to avoid obvious brightness abrupt changes or artificial artifacts. After reflection suppression, edge-preserving smoothing or median filtering can be selectively performed on the processed image to further suppress local noise while maintaining the integrity of the mucosal texture and edge structure. Finally, all processed images are recombined in the order of the original images to form a standard endoscopic image sequence that is consistent in terms of brightness, contrast, and reflective interference. The effective imaging mask and reflective area mask are also output as auxiliary data for subsequent steps of constructing anatomical constraints and analyzing abnormal textures.

[0018] Furthermore, reading the image sequence acquired during the continuous advancement of the endoscope also includes: Perform imaging state verification detection on the image sequence and establish verification detection results; if the verification detection result is a failure result, generate a re-acquisition instruction and perform re-acquisition processing on the image sequence according to the re-acquisition instruction.

[0019] Preferably, after acquiring the image sequence during the continuous advancement of the endoscope, the system performs an imaging state verification test on the image sequence to determine whether the current image sequence meets the imaging quality requirements for subsequent intelligent recognition and annotation. Specifically, firstly, basic quality indicators are calculated for the image sequence, including calculating the average brightness, brightness contrast, sharpness index, and effective imaging area ratio for each frame. The sharpness index can be characterized by gradient energy, Laplacian operator response, or the proportion of high-frequency energy in the frequency domain to determine if there is significant defocusing or motion blur. The effective imaging area ratio is calculated by the proportion of non-black-edge pixels in the effective field of view mask to detect lens occlusion or field of view shift. Subsequently, the proportion of specular reflection area in each frame of the image sequence is statistically analyzed. Based on the reflection area mask, the proportion of reflective pixels within the effective imaging area is calculated to determine if there is large-area strong reflection affecting lesion observation. Simultaneously, based on the brightness variation amplitude and the field of view center position variation amplitude between adjacent frames, the imaging stability of the image sequence is evaluated to detect abnormal imaging states such as severe shaking, rapid advancement, or abnormal regression. After calculating the above indicators, the brightness, sharpness, reflectivity, and imaging stability indicators are compared with preset imaging quality thresholds. If any indicator in the image sequence is below the corresponding threshold for multiple consecutive frames, or if the overall quality score (weighted result) is below the preset passing threshold, the current image sequence is determined to have failed the imaging status verification, and a corresponding verification detection result is generated; otherwise, the imaging status verification is determined to have passed.

[0020] For example, for an image with a resolution of 1280×720, it has a total of 921,600 pixels. After removing black borders, the effective imaging area is approximately 80%, with an effective pixel count of approximately 737,280. The detection frame count is 10 consecutive frames. When calculating the brightness index, the RGB image is first converted to the brightness channel Y using Y=0.299R+0.587G+0.114B. Then, the average brightness is calculated within the effective imaging area. The average brightness of the first frame F1 is 118, the average brightness of the second frame F2 is 121, and so on. These average brightness values ​​are all within the acceptable brightness range [110, 140]. When calculating the sharpness index, Laplacian filtering is first applied to the grayscale image, and then the variance of the Laplacian response is calculated. The Laplacian variance of F1 is 145, the Laplacian variance of F2 is 152, and so on. These Laplacian variances are all greater than the empirical sharpness threshold of 120. When calculating the reflectivity ratio, the reflectivity criteria are first defined as a brightness greater than or equal to 245 and a saturation less than or equal to 120. Then, the number of reflective pixels is counted based on these criteria, and by dividing this count by the total number of pixels, the reflectivity ratio for F1 is 2.44%, for F2 it is 2.85%, and so on. All of these reflectivity ratios are less than the reflectivity ratio threshold of 8%. When calculating the image stability index, the geometric center offset of the effective area between adjacent frames is calculated using Euclidean distance, and the brightness change rate of the effective area between adjacent frames is calculated using the absolute difference. The offset from F1 to F2 is 3.1, with a change of 3; the offset from F2 to F3 is 2.7, with a change of 2, and so on. All these offsets are less than the stability threshold of 10 pixels, and the change values ​​are all less than 15.

[0021] When the verification test result is unsuccessful, the system generates a re-acquisition instruction based on the specific imaging quality type of the failure. This re-acquisition instruction includes at least a re-acquisition trigger flag, a suggested re-acquisition timing, and corresponding imaging adjustment parameters. The imaging adjustment parameters may include prompts for the operator to slow down the advance speed, adjust the light source brightness, change the observation angle, or return to the previous effective field of view. These parameters are obtained by matching the deviation of the failed indicator with a control sample library. Finally, the re-acquisition instruction is sent to the endoscope host or display terminal to prompt the operator to perform image sequence re-acquisition processing until the re-acquisitioned image sequence passes the imaging status verification test. This ensures that subsequent intelligent recognition and image annotation processes are based on qualified and stable imaging data, improving the reliability and stability of lesion image annotation results.

[0022] Based on the cavity edge contour, mucosal fold direction, and tissue texture periodicity, a set of structural features is extracted to characterize the distribution of anatomical structures within the endoscopic field of view, and an anatomical structure constraint map is generated based on the set of structural features.

[0023] In one embodiment, after obtaining the standard endoscopic image sequence, a set of structural features is first constructed for each frame of the standard image to characterize the distribution of anatomical structures. This set of structural features includes cavity edge contours, mucosal fold orientations, and tissue texture periodicity features. The cavity edge contours are determined by multi-scale edge extraction, connected component analysis, and curve tracking within the effective field of view of a single frame image, and are used to exclude non-anatomical areas such as black borders and instrument obstructions. The mucosal fold orientations are determined by calculating the principal gradient direction and orientation concentration of candidate anatomical regions, as well as judging orientation consistency. The tissue texture periodicity features are determined by performing one-dimensional projection sampling of the pixel grayscale sequence along the principal gradient direction of each block, and then performing peak detection on the projection results. Finally, the set of pixels that simultaneously meet the criteria of being located in the candidate anatomical region, belonging to the mucosal fold direction region, and being determined to be a region with stable texture periodicity is marked as the effective anatomical structure region, and the remaining pixels are marked as the ineffective anatomical structure region. The effective anatomical structure region is assigned a value of 1, and the ineffective region is assigned a value of 0, generating a binary anatomical structure constraint map of the same size as the original image. This anatomical structure constraint map is used to limit the spatial range of subsequent abnormal texture analysis and lesion saliency calculation, thereby reducing the interference of normal fold texture on lesion candidate extraction and improving annotation stability.

[0024] Furthermore, extracting a set of structural features to characterize the distribution of anatomical structures within the endoscopic field of view, and generating an anatomical constraint map based on the set of structural features, includes: Multi-scale edge detection is performed on single frames of standard endoscopic image sequences. At least one continuous closed cavity wall edge curve is selected based on the closure and curvature change thresholds of the edge connectivity region. The region enclosed by these cavity wall edge curves is defined as a candidate anatomical region. Within the candidate anatomical region, the image pixel gradient directions are statistically analyzed in blocks, and the principal gradient direction and direction concentration parameter are calculated for each block. Blocks with a direction concentration parameter greater than a preset concentration threshold are determined as directionally reliable blocks, and direction consistency analysis is performed only on these directionally reliable blocks. Adjacent directionally reliable blocks are weighted based on the angle between the principal gradient directions and the direction concentration parameter. The directional consistency threshold is compared, and adjacent directional reliable blocks that meet the directional consistency threshold are merged to form the mucosal fold direction region. The pixel gray values ​​of the mucosal fold direction region are projected one-dimensionally along the main gradient direction, and the gray value repetition peak spacing and peak spacing variance are calculated using the projection signal. The texture periodic stable region is determined based on the calculation results. The set of pixels that are simultaneously located in the candidate anatomical region and the mucosal fold direction region and are determined to be the texture periodic stable region is marked as the anatomical structure effective region, and the remaining pixels are marked as the anatomical structure ineffective region. A binary anatomical structure constraint map is generated based on the anatomical structure effective region and the anatomical structure ineffective region.

[0025] Preferably, multi-scale edge detection processing is first performed on a single frame image within the effective field of view. Specifically, after Gaussian smoothing of the image at different scale parameters, the gradient magnitude and direction are calculated to obtain multi-scale edge response maps. Then, non-maximum suppression and thresholding are performed on each scale edge response map, and the multi-scale edge results are fused to obtain a stable set of edge candidates. After obtaining the edge candidate set, it is binarized. Pixels with edge responses greater than a preset threshold are marked as edge pixels, and the remaining pixels are marked as non-edge pixels. Connectivity analysis is then performed on the binary edge map, clustering the edge pixels according to their 8-neighborhood connections to obtain several independent edge connected regions. For each edge connected region, its boundary pixel is selected as the starting point, and adjacent edge pixels are traversed sequentially along the pixel connection direction. The pixels are sorted according to spatial continuity rules, thus transforming the disordered set of edge pixels into a sequentially arranged sequence of edge points, enabling edge curve tracking and reconstruction. During curve tracking, if the spatial distance or direction change between adjacent edge pixels exceeds a preset threshold, the curve is considered interrupted, and the current tracking is terminated. Through the above connected component analysis and curve tracing processing, the discrete and disordered edge responses in the edge candidate set can be organized into one or more continuous edge curves with clear geometric structures. Then, based on the closure of the edge connected region and whether the curvature change along the curve direction is continuous and lower than the preset curvature change threshold, at least one continuous closed edge curve that meets the conditions is retained as the cavity inner wall edge curve. The area enclosed by the cavity inner wall edge curve is defined as the candidate anatomical area, which is used to exclude non-anatomical areas such as black edges, bubbles, and instrument obstructions.

[0026] Within the candidate anatomical region, the region is divided into several local blocks of fixed or adaptive size. Histogram statistics are performed on the pixel gradient directions within each block to calculate the directional energy distribution and determine the dominant gradient direction of that block. Simultaneously, the proportion of dominant directional energy to total directional energy or the inverse of directional entropy within the block is calculated as a directional concentration parameter to measure the consistency of texture directions within the block. Blocks with a directional concentration parameter greater than a preset concentration threshold are then classified as directionally reliable blocks, and directional consistency analysis is performed only on these reliable blocks to avoid interference from directionally chaotic regions in subsequent wrinkle modeling. During this directional consistency analysis, the system calculates the angle difference between the dominant gradient directions of spatially adjacent directionally reliable blocks and, combined with the directional concentration parameters of adjacent blocks, generates a weighted directional consistency threshold. When the angle difference is less than the corresponding directional consistency threshold, adjacent directionally reliable blocks are considered to have directional consistency and are merged into the same directional region. Through this iterative merging process, several continuously distributed mucosal wrinkle directional regions are formed to describe the overall orientation structure of mucosal wrinkles within the endoscopic field of view. Within these mucosal fold regions, the system performs one-dimensional projection sampling of pixel grayscale values ​​along the principal gradient direction corresponding to each region. Specifically, the projection direction is determined based on the principal gradient direction corresponding to the region, and a set of equally spaced sampling lines are constructed along this direction to sample pixels point-by-point. For each sampling line, the grayscale values ​​of pixels located at the same spatial offset are accumulated or averaged, thereby compressing the grayscale distribution within the two-dimensional region into a one-dimensional sequence. Each sampling point in this one-dimensional sequence corresponds to a spatial position along the principal gradient direction, and its value represents the comprehensive grayscale response of pixels in the region at that position, thus obtaining a one-dimensional projection signal in which grayscale varies with spatial position. Then, peak detection is performed on the projection signal, the spacing between adjacent repeated peaks is calculated, and the mean and variance of the peak spacing are statistically analyzed. When the peak spacing is within a preset reasonable range and the peak spacing variance is less than a preset threshold, the corresponding region is determined to be a region with stable texture period, used to characterize the periodic texture structure of normal mucosal tissue. Finally, the set of pixels that simultaneously meet the criteria of being located in the candidate anatomical region, belonging to the mucosal fold direction region, and being determined to be a region with stable texture periodicity is marked as the effective anatomical structure region, and the remaining pixels are marked as the ineffective anatomical structure region. Then, a binary anatomical structure constraint map with the same size as the original image is constructed based on the effective anatomical structure region and the ineffective anatomical structure region. The effective anatomical structure region is assigned a first preset value, and the ineffective anatomical structure region is assigned a second preset value. This anatomical structure constraint map is used to limit the spatial range of subsequent abnormal texture perturbation analysis and lesion saliency candidate extraction, thereby improving the accuracy and stability of lesion image annotation.

[0027] For example, during a gastroscopy, a standard endoscopic image with a resolution of 1280×1024 is acquired as the endoscope continuously advances through the gastric body. Multi-scale edge detection is performed on this image, extracting edge responses at three scales: σ=1.0, σ=2.0, and σ=3.0, and fusing them to obtain a candidate edge set. After connected component analysis, a closed edge curve with a length of approximately 3120 pixels is detected in the central region of the image. The average curvature change rate along the curve is 0.018, which is less than the preset curvature threshold of 0.03. Therefore, this closed curve is identified as the edge curve of the gastric cavity wall, and the approximately 82% of the image area enclosed by it is defined as the candidate anatomical region. Within the candidate anatomical region, the image was divided into 32×32 pixel blocks, resulting in approximately 1000 blocks. After analyzing the pixel gradient direction histogram for each block, it was found that approximately 720 blocks had a principal gradient direction energy ratio greater than 0.65, exceeding the direction concentration threshold of 0.6, and were therefore identified as directionally reliable blocks. Next, direction consistency analysis was performed on spatially adjacent directionally reliable blocks. One pair of adjacent blocks had principal gradient directions of 42° and 48°, with an included angle of 6°. After weighting the direction concentration parameter, the direction consistency threshold was 10°, satisfying the consistency condition, and therefore, they were merged into the same mucosal fold direction region. Through this merging process, five main mucosal fold direction regions were ultimately formed. Within one of the mucosal fold directions, a one-dimensional projection of the pixel grayscale values ​​is performed along its principal gradient direction (approximately 45°), resulting in a grayscale variation sequence of length 256. Peak detection then identifies seven recurring grayscale peaks with an average spacing of 18.4 pixels and a variance of 2.1, all within the preset periodicity range (15–25 pixels, variance less than 5). Therefore, this region is classified as a texture periodicity stable region. Finally, the set of pixels that simultaneously satisfy the criteria of candidate anatomical region, mucosal fold direction region, and texture periodicity stable region is marked as the effective anatomical structure region. A binary anatomical structure constraint map with the same size as the original image is generated. The effective region pixels account for approximately 68% of the entire frame image, while the remaining regions are marked as ineffective anatomical structure regions for subsequent abnormal texture perturbation analysis.

[0028] Within the spatial range defined by the anatomical structure constraint map, local directional texture consistency analysis of the standard endoscopic image sequence is performed to establish abnormal texture perturbation components.

[0029] In one embodiment, after obtaining the anatomical structure constraint map, the system uses pixels marked as valid anatomical structure areas in the constraint map as the analysis objects, limiting subsequent operations to be performed only within this spatial range to eliminate interference from black borders, reflective residues, and non-valid tissue areas on texture judgment. Subsequently, for a single frame image in the standard endoscopic image sequence, the system first divides the valid anatomical structure area into local blocks of a preset block size (e.g., 16×16 or 32×32 pixels), then calculates the pixel gradient direction histogram for each block and determines the main texture direction of that block. Afterwards, using the main texture direction as the axis, one-dimensional grayscale sampling is performed within the block along the main texture direction and in a direction perpendicular to the main texture direction, respectively, to obtain grayscale change sequences associated with the two directions. Based on these two grayscale change sequences, the system calculates the consistency index of the grayscale changes, compares the consistency index along the main texture direction and the perpendicular direction, and calculates the difference as the texture direction consistency judgment metric. When the judgment value exceeds a preset threshold, it indicates that the grayscale variation pattern of the block deviates abnormally in the main texture direction and vertical direction, exhibiting perturbation characteristics that do not conform to the normal mucosal fold texture. In this case, the judgment value is mapped to an abnormal texture perturbation component, and the abnormal texture perturbation component can be further backfilled to the corresponding pixel by combining the pixel position within the block, forming an abnormal perturbation response map of the same size as the original image. Through the above local directional texture consistency analysis within the range of the anatomical structure constraint map, the directional texture of normal mucosal folds can be used as a background benchmark, and the local texture perturbation that deviates from the benchmark can be highlighted in the form of numerical components, providing stable input for subsequent multi-scale aggregation to generate significant lesion candidate maps.

[0030] Furthermore, within the spatial range defined by the anatomical constraint map, local orientation texture consistency analysis of the endoscopic standard image sequence is performed to establish anomalous texture perturbation components, including: Within the spatial range defined by the anatomical structure constraint diagram, a single frame image of the endoscopic standard image sequence is partially segmented. The gradient direction of pixels within each segment is statistically analyzed to determine the main texture direction of the segment. One-dimensional sampling is performed on the grayscale values ​​of pixels within the segment along the main texture direction and the direction perpendicular to the main texture direction to obtain a direction-related grayscale change sequence. The grayscale change consistency index along the main texture direction and the direction perpendicular to the main texture direction is calculated using the grayscale change sequence, and the difference between the grayscale change consistency indices is used as the texture direction consistency judgment quantity. Abnormal texture perturbation components are configured using the texture direction consistency judgment quantity.

[0031] Preferably, within the spatial range defined by the anatomical structure constraint map, local orientation texture consistency analysis is performed on single-frame images in the standard endoscopic image sequence only for pixels marked as valid anatomical structure regions. Specifically, firstly, the single-frame image is locally divided into blocks according to a preset size within the spatial range. The block size can be set to a fixed size or an adaptive size according to the imaging resolution, and adjacent blocks may overlap. After the division, the pixel gradient direction within each block is statistically analyzed to construct a gradient direction histogram, and the direction with the largest energy proportion is selected as the main texture direction of the block, used to characterize the main direction of the mucosal texture within the block. After determining the main texture direction of the block, one-dimensional sampling is performed on the pixel grayscale values ​​within the block along the main texture direction and in the direction perpendicular to the main texture direction, respectively, to obtain two grayscale change sequences associated with the direction. The grayscale change sequence consists of grayscale values ​​sampled at equal intervals along the corresponding direction within the block, used to reflect the law of grayscale change with spatial position. Subsequently, grayscale variation consistency indices are calculated based on two grayscale variation sequences. Specifically, the grayscale variation sequences are differentially processed to calculate the grayscale variation between adjacent sampling points, resulting in a grayscale variation difference sequence. This difference sequence reflects the fluctuation range of grayscale in the spatial direction. Consistency measurement parameters, including the variance, root mean square error, or absolute deviation of the difference sequence, are then calculated based on this difference sequence. A smaller variance or root mean square error indicates stronger consistency and stability of grayscale variation in that direction; conversely, a larger variance or root mean square error indicates significant fluctuations and weaker regularity in grayscale variation. The calculated grayscale variation consistency index along the main texture direction is used to characterize the texture stability along the normal mucosal fold direction, while the grayscale variation consistency index perpendicular to the main texture direction is used to characterize the texture variation characteristics across fold directions. Finally, the difference between the grayscale variation consistency index calculated along the main texture direction and the grayscale variation consistency index calculated perpendicular to the main texture direction is calculated, and this difference is used as the texture direction consistency criterion. When the texture direction consistency determination value exceeds the preset threshold, it indicates that there is a significant difference in the consistency of grayscale changes in the main texture direction and the vertical direction of the block, and there are abnormal texture perturbation features that deviate from the normal mucosal texture structure. At this time, the texture direction consistency determination value is mapped to the abnormal texture perturbation component, and the abnormal texture perturbation component is associated with the pixel position in the block to construct the subsequent abnormal texture perturbation response map, thereby providing basic input for multi-scale aggregation and lesion saliency candidate generation.

[0032] Multi-scale aggregation is performed on the abnormal texture perturbation components to establish a significant candidate map of lesions. Based on the standard endoscopic image sequence, the spatial position continuity analysis of adjacent frames of the significant candidate map of lesions is performed to calculate the position offset and morphological change, and to construct a continuous correction result.

[0033] In one embodiment, after obtaining the anomalous texture perturbation component of a single-frame image, the system performs multi-scale aggregation on the perturbation component. During this process, at the pixel scale, the anomalous texture perturbation component is backfilled into a perturbation response map of the same size as the original image. The perturbation energy of each pixel is calculated and normalized within a preset neighborhood to obtain a pixel-level anomalous significant response. Subsequently, at the local region scale, the image is divided into overlapping blocks with a fixed step size. Within each block, directional consistency aggregation and energy normalization are performed on the perturbation response to obtain a local region-level anomalous significant response, used to suppress isolated noise points and enhance regional anomalies. Next, at the structural block scale, adjacent blocks are combined into structural block-level analysis units based on the spatial topological relationships of the anatomical structure constraint map. The cross-regional cooperative variability of the perturbation response is calculated within the structural block to form a structural block-level anomalous significant response, highlighting anomalous regions that are continuously distributed within the effective area of ​​the anatomical structure. Subsequently, the salient responses at the pixel level, local region level, and structural block level are weighted and fused to obtain a salient candidate map of lesions. Threshold segmentation and connected component filtering are then used to extract candidate salient regions and their attributes, such as boundary contours, geometric centers, and areas. After generating the salient candidate map, the system uses the geometric centers, boundary contour shape descriptors, and area distributions of candidate salient regions in adjacent frames as joint spatial descriptors for cross-frame matching to determine corresponding salient region pairs. For the matched salient region pairs, the system calculates the positional offset, area change rate, perimeter change rate, or contour difference as morphological change parameters. The offsets and changes across multiple consecutive frames are then smoothly estimated to obtain continuous correction results. When a sudden jump or abnormal morphological fluctuation is detected in a single-frame candidate region, the continuous correction results are used to correct the position and boundary of the candidate region in that frame. This ensures that the spatial position and morphological changes of the lesion candidates maintain reasonable continuity during continuous advancement, thereby reducing candidate drift and false detections caused by jitter and viewpoint changes, and providing stable lesion candidate input for subsequent image annotation and report generation.

[0034] Furthermore, multi-scale aggregation is performed on the abnormal texture perturbation components to establish a significant candidate map of lesions, including: Based on the endoscopic imaging resolution and anatomical structure response map, a pixel-level response layer with single-pixel neighborhoods as basic units is constructed on the original image size. Within the pixel-level response layer, the local directional energy distribution of abnormal texture perturbation components within a preset direction set is calculated to generate a pixel-level abnormal saliency response map. Based on the pixel-level abnormal saliency response map, a single frame image is overlapped and divided into blocks according to a fixed spatial step size to construct a local region-level response layer. Within each local region, directional consistency aggregation and regional energy normalization processing are performed on the abnormal texture perturbation components to generate a local region-level abnormal saliency response map. On the local region-level response layer, based on the spatial topological relationship defined by the anatomical structure constraint map, multiple adjacent local regions are combined into a structural block-level analysis unit. Within the structural block-level analysis unit, the cross-regional cooperative variation degree of abnormal texture perturbation components is calculated to generate a structural block-level abnormal saliency response map. Aggregate analysis of the pixel-level abnormal saliency response map, the local region-level abnormal saliency response map, and the structural block-level abnormal saliency response map is performed to establish a significant candidate map of lesions.

[0035] Preferably, after obtaining the anomalous texture perturbation components corresponding to a single frame image, a pixel-level response layer is first constructed on the original image size. That is, using a single pixel and its neighboring pixels as the basic analysis unit, directional energy statistics are performed on the anomalous texture perturbation components corresponding to each pixel within the range marked as the effective region of the anatomical structure constraint map. Specifically, in a preset set of directions, the amplitudes of the anomalous texture perturbation components in the neighborhood are accumulated to obtain the local directional energy distribution corresponding to each direction. The directional energy is then normalized, and the energy value corresponding to the dominant direction in the directional energy distribution is taken as the anomalous saliency response intensity of that pixel, thereby generating a pixel-level anomalous saliency response map with the same size as the original image, used to characterize fine-grained, localized anomalous texture responses. Subsequently, using the pixel-level anomalous saliency response map as input, the single frame image is processed by overlapping block segmentation according to a fixed spatial step size, for example, with 32×32 pixels as a block and a step size of 16 pixels. Within each local region, statistical analysis is performed on the anomalous salient responses of all pixels within the region. The consistency of the anomalous salient responses in each direction is calculated to determine whether the anomalous responses within the region are spatially concentrated. For regions with high directional consistency, their anomalous salient responses are weighted and accumulated, while high-response pixels that are directionally dispersed or isolated are suppressed. Simultaneously, the aggregated anomalous energy within the region is normalized to make the response intensities between different regions comparable. Finally, a local region-level anomalous salient response map is generated to highlight anomalous regions with a certain area and continuity. Based on the local region-level response layer, the system combines multiple spatially adjacent local regions belonging to the effective region of the anatomical structure into structural block-level analysis units according to the spatial topological relationships defined by the anatomical structure constraint map. Within each structural block-level analysis unit, the degree of cooperative variation of the anomalous texture perturbation components between local regions is calculated. This degree of cooperative variation characterizes the consistency of the anomalous salient responses of adjacent local regions in terms of intensity variation trend, directional distribution, and spatial location. It is obtained by averaging the inter-regional consistency scores of all adjacent local region pairs. When the anomalous responses of multiple local regions exhibit synchronous enhancement in space, the saliency response weight of the corresponding structural blocks is increased, thereby generating a structural block-level anomalous saliency response map. This map is used to enhance anomalous regions that conform to anatomical structural continuity and suppress scattered noise. Finally, aggregation analysis is performed on the pixel-level, local region-level, and structural block-level anomalous saliency response maps. The three types of response maps are superimposed using a weighted fusion method to obtain a lesion saliency candidate map that comprehensively reflects local detail anomalies, regional continuity anomalies, and structural consistency anomalies. The high-response areas in this lesion saliency candidate map are potential lesion candidate regions, providing stable and reliable input for subsequent cross-frame continuity analysis, location correction, and image annotation.

[0036] Furthermore, constructing significant candidate maps of lesions includes: For the salient response maps at each scale, the response stability in adjacent scales and adjacent spatial units is calculated, and the spatial stability weights corresponding to the scales are constructed using the response stability. The consistency of the abnormal texture perturbation components in the dominant direction is calculated in the salient response maps at each scale, and texture consistency weights are constructed. Multi-scale weighted aggregation and fusion are performed using the spatial stability weights and texture consistency weights to establish salient candidate maps of lesions.

[0037] Optionally, after obtaining pixel-level, local region-level, and structural block-level saliency response maps, respectively, the response stability of each scale's saliency response map in adjacent scales and adjacent spatial units is calculated to construct spatial stability weights corresponding to each scale. Specifically, within the same spatial location or corresponding spatial unit, the difference between the current scale response value and the response values ​​of adjacent scales is compared. When the response values ​​at different scales are consistent or change gradually in amplitude and spatial distribution, the location or unit is determined to have high scale stability. Specifically, the scale response value at the pixel level is obtained by statistically analyzing and normalizing the local directional energy of the anomalous texture perturbation components in the preset direction set of the pixel and its neighborhood; the scale response value at the local region level is obtained by weighted accumulation after filtering the pixel-level saliency responses of all pixels in the region for directional consistency, and the accumulation result is normalized according to the region area or the number of effective pixels; the scale response value at the structural block level is obtained by weighted fusion of the regional scale response values ​​of each local region. Simultaneously, within the same scale, the anomalous saliency responses of spatially adjacent pixels, blocks, or structural blocks are compared and analyzed. When the response intensity and trend of adjacent spatial units are consistent, the region is considered to have high spatial stability. Scale stability and spatial stability are combined to obtain the spatial stability weight corresponding to each scale. This weight reflects the reliability of the anomalous response at that scale in both spatial and scale dimensions. Then, in each scale response map, the dominant texture direction of the corresponding region or unit is determined according to the texture analysis steps, and the distribution continuity and directional concentration of the anomalous saliency responses are statistically analyzed along this dominant direction. When the anomalous response shows a continuous distribution in the dominant direction and maintains a similar intensity trend in adjacent spatial locations, its texture consistency is considered to be high; conversely, if the anomalous response shows a broken or dispersed characteristic in the dominant direction, its texture consistency is considered low. Then, based on the consistency retention degree, corresponding texture consistency weights are constructed for each scale response to suppress isolated anomalous responses that do not match the dominant texture direction. Finally, for the anomalous saliency responses from different scales at the same spatial location, they are multiplied by the corresponding scale's spatial stability weight and texture consistency weight, and then weighted and summed to obtain the comprehensive anomalous saliency response value. Through the above fusion process, abnormal regions that are stable and consistent in texture across multiple scales and spatial units will be significantly enhanced, while abnormal responses that are unstable in scale and inconsistent in direction will be suppressed. Finally, a significant candidate map of lesions will be established. This significant candidate map of lesions can simultaneously reflect the local detail features and regional continuity features of abnormal textures, providing a highly reliable input basis for subsequent cross-frame continuity analysis, candidate region selection and image annotation.

[0038] Furthermore, based on standard endoscopic image sequences, spatial continuity analysis of adjacent frames of salient candidate lesion maps is performed, including: Between two adjacent frames of standard endoscopic images, the geometric center, boundary contour, and area distribution of salient regions in the salient candidate images of lesions are used as joint spatial descriptors to perform cross-frame region matching and calculate the positional offset and morphological change of the corresponding salient regions.

[0039] Preferably, after obtaining salient candidate images of lesions corresponding to two adjacent frames of standard endoscopic images, the system uses the salient regions extracted from the candidate images as the analysis objects. First, it performs threshold segmentation and connected component analysis on the salient candidate images of each frame to extract one or more salient regions. Geometric feature parameters are then calculated for each salient region. These geometric feature parameters include at least the geometric center coordinates, boundary contours, and area distribution information represented by the number of pixels, used to construct a region-level joint spatial descriptor. Subsequently, using each salient region in the current frame as a reference, the system searches for candidate regions in the set of salient regions in the next frame that are spatially closest and morphologically similar. Specifically, coarse matching is performed based on the Euclidean distance between geometric centers, filtering out candidate regions whose distance is less than a preset search radius. Based on this, the similarity of boundary contours and area distribution is further compared among the candidate regions. Boundary contour similarity can be evaluated by contour overlap, shape similarity measurement, or difference in the principal axis direction of the contours, while area distribution similarity can be evaluated by area ratio or area change rate. By comprehensively considering the geometric center distance, contour similarity, and area similarity, the system can select the pair of salient regions with the highest similarity in the joint spatial descriptor as the corresponding regions across frames. After identifying the corresponding salient regions, the positional offset and morphological change of the corresponding salient regions in two adjacent frames are calculated. The positional offset is obtained by calculating the Euclidean distance between the geometric center coordinates of the salient regions in the two frames, which is used to characterize the spatial movement of the lesion during endoscopic advancement. The morphological change is obtained by calculating the area change rate, perimeter change rate, or contour difference of the salient regions in the two frames, which is used to characterize the changes in the morphology of the salient regions with changes in viewing angle or advancement depth. The area change rate can be calculated by the ratio of the area difference of the corresponding salient regions in the two adjacent frames to the area of ​​the salient region in the previous frame. The perimeter change rate can be calculated by the ratio of the perimeter difference of the corresponding salient regions in the two adjacent frames to the perimeter of the salient region in the previous frame. The contour difference can be calculated by aligning the boundary contours of the corresponding salient regions in the two adjacent frames and then calculating the average distance between the contour points. Finally, the positional offset and morphological change are used as the quantitative results of cross-frame continuity analysis, providing a basis for subsequent continuous correction, abnormal jump suppression and image annotation stability enhancement. This enables reliable correspondence between lesion candidate regions in adjacent frames under continuous endoscopic advancement and viewing angle changes, suppressing mismatches and jumps caused by imaging jitter or changes in field of view, and improving the continuity and stability of image annotation.

[0040] Image annotation of the endoscopic standard image sequence is performed using the continuous correction results.

[0041] In one embodiment, after obtaining the cross-frame continuity analysis results of significant candidate lesion images, the system constructs continuous correction results based on positional offset and morphological changes, and uses these continuous correction results to perform image annotation on the standard endoscopic image sequence. Specifically, for the candidate lesion regions extracted from each frame image, their spatial positions are first corrected based on the continuous correction results. When the geometric center of a candidate region in a certain frame experiences a sudden displacement exceeding a preset threshold relative to the preceding and following frames, the geometric center of that frame is regressed and corrected using the displacement trend of the corresponding regions in the preceding and following frames, and the boundary of the candidate region in that frame is translated accordingly. Simultaneously, the morphology of the candidate region is corrected; that is, when the morphological changes such as the rate of change of area and the degree of difference in contour exceed a reasonable range, the contour of the candidate region in that frame is smoothed and repaired by combining the principal axis direction of the contour of the adjacent frame regions and the average area, in order to suppress false abnormal regions caused by reflection, jitter, and sudden changes in viewpoint. After correction, annotation information is output based on the corrected candidate regions. This includes generating corresponding bounding boxes or pixel-level masks on the original image and assigning unique identifiers to the annotated objects for cross-frame tracking. Furthermore, the location and size parameters of lesions, such as equivalent diameter, major and minor axis lengths, and area, are calculated by combining the geometric center, contour boundaries, and pixel area of ​​the candidate regions. Color features, such as average hue, red composite index, or brightness distribution, are statistically analyzed within the candidate regions to form quantitative descriptive fields of what is seen under the microscope. Finally, the above-mentioned location, morphology, and color annotation fields are associated with frame sequence numbers and timestamps to generate structured annotation results. This achieves continuous and consistent image annotation of standard endoscopic image sequences, providing stable input for subsequent selection of keyframes with clinical reference value and automatic generation of image and text reports.

[0042] Furthermore, performing image annotation of the endoscopic standard image sequence using the continuous correction results includes: A graphic report is generated based on the image annotation results, and the graphic report is sent to the display terminal for visual interactive annotation display.

[0043] Preferably, after completing the image annotation of the standard endoscopic image sequence, the system automatically generates a graphic report based on the image annotation results and sends the graphic report to the display terminal for visual interactive annotation display. Specifically, firstly, the annotation objects corresponding to the cross-frames are aggregated according to the lesion identifier, and candidate lesion regions that are stably present in multiple consecutive frames and whose significance is higher than a preset threshold are selected as key annotation objects with clinical reference value. For each key annotation object, the representative frame with the highest significance or the best imaging quality is selected as the key frame image, and its occurrence interval information in the time series is retained. Subsequently, at the image level, the lesion annotation results are superimposed on the key frame image, including the lesion boundary contour, center mark, number, etc.; at the text level, the description field of the endoscopic view is automatically generated based on the structured parameters of the annotation object, including at least the location description of the lesion in the endoscopic field of view, geometric size parameters, morphological features, color features, and the number of consecutive frames. By arranging the image information and text description according to the preset report template, a structured graphic report can be formed. After the graphic report is generated, it is sent to a display terminal for visualization. This terminal displays keyframe images and their corresponding textual descriptions, supporting interactive browsing based on a timeline or lesion list. Operators can zoom in and out, switch frames, hide or show annotation layers, and confirm, modify, or supplement the automatically generated annotations on the display terminal. This interactive process is recorded and fed back to the system to update the final report, thus achieving intuitive display and efficient clinical application of endoscopic image annotation results.

[0044] In summary, the embodiments of this application have at least the following technical effects: This method involves reading image sequences acquired during continuous endoscopy, performing brightness normalization and specular reflection suppression on each frame of the sequence to generate a standard endoscopic image sequence with consistent imaging conditions. Based on cavity edge contours, mucosal fold orientation, and tissue texture periodicity, a set of structural features characterizing the distribution of anatomical structures within the endoscopic field of view is extracted, and an anatomical constraint map is generated. Within the spatial range defined by the anatomical constraint map, local directional texture consistency analysis of the standard endoscopic image sequence is performed to establish abnormal texture perturbation components. These abnormal texture perturbation components are then aggregated at multiple scales to establish significant lesion candidate images. Based on the standard endoscopic image sequence, spatial position continuity analysis of adjacent frames of the significant lesion candidate images is performed, calculating positional offset and morphological changes to construct continuous correction results. Finally, image annotation of the standard endoscopic image sequence is performed using these continuous correction results. This method solves the technical problems of unstable image quality and low image annotation accuracy during continuous endoscopic image acquisition, achieving the technical effect of improving image quality consistency, stability, and annotation accuracy by fusing anatomical prior constraints and multi-frame dynamic continuity analysis.

[0045] Example 2, based on the same inventive concept as the intelligent recognition and annotation method for endoscopic images in the foregoing examples, such as... Figure 2 As shown, this application provides an intelligent recognition and annotation system for endoscopic images, wherein the system includes: Image standardization module 11: Reads the image sequence acquired during continuous endoscopy, performs brightness normalization and specular reflection suppression processing on each frame of the image sequence, and generates a standard endoscopic image sequence with consistent imaging conditions; Feature extraction module 12: Extracts a set of structural features to characterize the distribution of anatomical structures within the endoscopic field of view based on cavity edge contours, mucosal fold direction, and tissue texture periodicity features, and generates an anatomical structure constraint map based on the set of structural features; Consistency analysis module 13: Performs local directional texture consistency analysis on the standard endoscopic image sequence within the spatial range defined by the anatomical structure constraint map, and establishes abnormal texture perturbation components; Continuity analysis module 14: Performs multi-scale aggregation on the abnormal texture perturbation components, establishes significant lesion candidate maps, and performs spatial position continuity analysis of adjacent frames of significant lesion candidate maps based on the standard endoscopic image sequence, calculates position offset and morphological change, and constructs continuous correction results; Image annotation module 15: Performs image annotation on the standard endoscopic image sequence using the continuous correction results.

[0046] Furthermore, the image normalization module 11 is used to perform the following method: Perform imaging state verification detection on the image sequence and establish verification detection results; if the verification detection result is a failure result, generate a re-acquisition instruction and perform re-acquisition processing on the image sequence according to the re-acquisition instruction.

[0047] Furthermore, the feature extraction module 12 is used to perform the following method: Multi-scale edge detection is performed on single frames of standard endoscopic image sequences. At least one continuous closed cavity wall edge curve is selected based on the closure and curvature change thresholds of the edge connectivity region. The region enclosed by these cavity wall edge curves is defined as a candidate anatomical region. Within the candidate anatomical region, the image pixel gradient directions are statistically analyzed in blocks, and the principal gradient direction and direction concentration parameter are calculated for each block. Blocks with a direction concentration parameter greater than a preset concentration threshold are determined as directionally reliable blocks, and direction consistency analysis is performed only on these directionally reliable blocks. Adjacent directionally reliable blocks are weighted based on the angle between the principal gradient directions and the direction concentration parameter. The directional consistency threshold is compared, and adjacent directional reliable blocks that meet the directional consistency threshold are merged to form the mucosal fold direction region. The pixel gray values ​​of the mucosal fold direction region are projected one-dimensionally along the main gradient direction, and the gray value repetition peak spacing and peak spacing variance are calculated using the projection signal. The texture periodic stable region is determined based on the calculation results. The set of pixels that are simultaneously located in the candidate anatomical region and the mucosal fold direction region and are determined to be the texture periodic stable region is marked as the anatomical structure effective region, and the remaining pixels are marked as the anatomical structure ineffective region. A binary anatomical structure constraint map is generated based on the anatomical structure effective region and the anatomical structure ineffective region.

[0048] Furthermore, the consistency analysis module 13 is used to perform the following methods: Within the spatial range defined by the anatomical structure constraint diagram, a single frame image of the endoscopic standard image sequence is partially segmented. The gradient direction of pixels within each segment is statistically analyzed to determine the main texture direction of the segment. One-dimensional sampling is performed on the grayscale values ​​of pixels within the segment along the main texture direction and the direction perpendicular to the main texture direction to obtain a direction-related grayscale change sequence. The grayscale change consistency index along the main texture direction and the direction perpendicular to the main texture direction is calculated using the grayscale change sequence, and the difference between the grayscale change consistency indices is used as the texture direction consistency judgment quantity. Abnormal texture perturbation components are configured using the texture direction consistency judgment quantity.

[0049] Furthermore, the continuity analysis module 14 is used to perform the following methods: Based on the endoscopic imaging resolution and anatomical structure response map, a pixel-level response layer with single-pixel neighborhoods as basic units is constructed on the original image size. Within the pixel-level response layer, the local directional energy distribution of abnormal texture perturbation components within a preset direction set is calculated to generate a pixel-level abnormal saliency response map. Based on the pixel-level abnormal saliency response map, a single frame image is overlapped and divided into blocks according to a fixed spatial step size to construct a local region-level response layer. Within each local region, directional consistency aggregation and regional energy normalization processing are performed on the abnormal texture perturbation components to generate a local region-level abnormal saliency response map. On the local region-level response layer, based on the spatial topological relationship defined by the anatomical structure constraint map, multiple adjacent local regions are combined into a structural block-level analysis unit. Within the structural block-level analysis unit, the cross-regional cooperative variation degree of abnormal texture perturbation components is calculated to generate a structural block-level abnormal saliency response map. Aggregate analysis of the pixel-level abnormal saliency response map, the local region-level abnormal saliency response map, and the structural block-level abnormal saliency response map is performed to establish a significant candidate map of lesions.

[0050] Furthermore, the continuity analysis module 14 is used to perform the following methods: For the salient response maps at each scale, the response stability in adjacent scales and adjacent spatial units is calculated, and the spatial stability weights corresponding to the scales are constructed using the response stability. The consistency of the abnormal texture perturbation components in the dominant direction is calculated in the salient response maps at each scale, and texture consistency weights are constructed. Multi-scale weighted aggregation and fusion are performed using the spatial stability weights and texture consistency weights to establish salient candidate maps of lesions.

[0051] Furthermore, the continuity analysis module 14 is used to perform the following methods: Between two adjacent frames of standard endoscopic images, the geometric center, boundary contour, and area distribution of salient regions in the salient candidate images of lesions are used as joint spatial descriptors to perform cross-frame region matching and calculate the positional offset and morphological change of the corresponding salient regions.

[0052] Furthermore, the image annotation module 15 is used to perform the following method: A graphic report is generated based on the image annotation results, and the graphic report is sent to the display terminal for visual interactive annotation display.

[0053] Example 3, Figure 3 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention, showing a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present invention. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality or scope of the embodiments of the present invention. Figure 3As shown, the electronic device includes a processor 21, a memory 22, an input device 23, and an output device 24; the number of processors 21 in the electronic device can be one or more. Figure 3 Taking a processor 21 as an example, the processor 21, memory 22, input device 23, and output device 24 in an electronic device can be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.

[0054] The memory 22, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the intelligent recognition and annotation method for endoscopic images in this embodiment of the invention. The processor 21 executes various functional applications and data processing of the electronic device by running the software programs, instructions, and modules stored in the memory 22, thereby realizing the aforementioned intelligent recognition and annotation method for endoscopic images.

[0055] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. An intelligent recognition and annotation method for endoscopic images, characterized in that, The method includes: Read the image sequence acquired during the continuous advancement of the endoscope, perform brightness normalization and specular reflection suppression processing on each frame of the image sequence, and generate a standard endoscope image sequence with consistent imaging conditions; Based on the cavity edge contour, mucosal fold direction, and tissue texture periodicity, a set of structural features is extracted to characterize the distribution of anatomical structures within the endoscopic field of view, and an anatomical structure constraint map is generated based on the set of structural features. Within the spatial range defined by the anatomical structure constraint map, perform local orientation texture consistency analysis on the standard endoscopic image sequence to establish abnormal texture perturbation components; Multi-scale aggregation is performed on the abnormal texture perturbation components to establish a significant candidate map of lesions. Based on the standard endoscopic image sequence, the spatial position continuity analysis of adjacent frames of the significant candidate map of lesions is performed to calculate the position offset and morphological change, and to construct a continuous correction result. Image annotation of the endoscopic standard image sequence is performed using the continuous correction results.

2. The intelligent recognition and annotation method for endoscopic images as described in claim 1, characterized in that, Extracting a set of structural features to characterize the distribution of anatomical structures within the endoscopic field of view, and generating an anatomical constraint map based on the set of structural features, including: Multi-scale edge detection is performed on single-frame images in the standard endoscopic image sequence. At least one continuous closed cavity wall edge curve is selected based on the closure and curvature change threshold of the edge connected region. The region enclosed by the cavity wall edge curve is defined as the candidate anatomical region. Within the candidate anatomical region, the gradient directions of image pixels are statistically analyzed in blocks, and the principal gradient direction and direction concentration parameters within each block are calculated. Blocks whose directional concentration parameter is greater than a preset concentration threshold are identified as directionally reliable blocks, and directional consistency analysis is performed only on directionally reliable blocks. Adjacent directional reliable blocks are compared with the directional consistency judgment threshold obtained by weighting the principal gradient direction angle and the directional consistency judgment threshold based on the directional concentration parameter. Adjacent directional reliable blocks that meet the directional consistency judgment threshold are merged to form mucosal fold directional regions. The pixel grayscale values ​​are projected one-dimensionally along the principal gradient direction in the mucosal fold direction region, and the grayscale repetition peak spacing and peak spacing variance are calculated using the projection signal. The texture periodic stable region is determined based on the calculation results. The set of pixels that are simultaneously located in the candidate anatomical region, the mucosal fold direction region and are determined to be a texture periodic stable region is marked as the effective region of the anatomical structure, and the remaining pixels are marked as the ineffective region of the anatomical structure. A binary anatomical structure constraint map is generated based on the effective and ineffective regions of the anatomical structure.

3. The intelligent recognition and annotation method for endoscopic images as described in claim 1, characterized in that, Within the spatial range defined by the anatomical constraint map, local orientation texture consistency analysis of the standard endoscopic image sequence is performed to establish anomalous texture perturbation components, including: Within the spatial range defined by the anatomical structure constraint map, the single-frame images of the standard endoscopic image sequence are partially segmented, and the gradient direction of the pixels in each segment is statistically analyzed to determine the main texture direction of the segment. One-dimensional sampling is performed on the grayscale values ​​of pixels within the block along the direction of the main texture of the block and the direction perpendicular to the direction of the main texture of the block to obtain the direction-related grayscale change sequence. The grayscale change consistency index is calculated using the grayscale change sequence along the main texture direction and in the direction perpendicular to the main texture direction, and the difference in the grayscale change consistency index is used as the texture direction consistency judgment quantity. The abnormal texture perturbation component is configured using the texture direction consistency determination quantity.

4. The intelligent recognition and annotation method for endoscopic images as described in claim 1, characterized in that, Multi-scale aggregation of the abnormal texture perturbation components is performed to establish a significant candidate map of lesions, including: Based on the endoscopic imaging resolution and anatomical structure response map, a pixel-level response layer with single-pixel neighborhood as the basic unit is constructed on the original image size. Within the pixel-level response layer, the local directional energy distribution of abnormal texture perturbation components in a preset direction set is calculated to generate a pixel-level abnormal significant response map. Based on the pixel-level salient response map, the single frame image is overlapped and divided into blocks according to a fixed spatial step size to construct a local region-level response layer. In each local region, the abnormal texture perturbation component is subjected to directional consistency aggregation and regional energy normalization processing to generate a local region-level salient response map. At the local region level response layer, based on the spatial topological relationship defined by the anatomical structure constraint map, multiple adjacent local regions are combined into a structural block level analysis unit, and the cross-regional cooperative variation degree of the abnormal texture perturbation component is calculated within the structural block level analysis unit to generate a structural block level abnormal saliency response map. Perform aggregated analysis of pixel-level, local region-level, and structural block-level anomaly response maps to establish significant candidate maps of lesions.

5. The intelligent recognition and annotation method for endoscopic images as described in claim 4, characterized in that, Establish significant candidate maps of lesions, including: For the anomalous saliency response maps at each scale, the response stability in adjacent scales and adjacent spatial units is calculated respectively, and the spatial stability weights corresponding to the scale are constructed using the response stability. The consistency of anomalous texture perturbation components in the dominant direction is calculated in the anomalous saliency response map at each scale, and texture consistency weights are constructed. Multi-scale weighted aggregation and fusion are performed using spatial stability weights and texture consistency weights to establish significant candidate maps of lesions.

6. The intelligent recognition and annotation method for endoscopic images as described in claim 1, characterized in that, Based on standard endoscopic image sequences, spatial continuity analysis of adjacent frames of salient candidate lesion maps is performed, including: Between two adjacent frames of standard endoscopic images, the geometric center, boundary contour, and area distribution of salient regions in the salient candidate images of lesions are used as joint spatial descriptors to perform cross-frame region matching and calculate the positional offset and morphological change of the corresponding salient regions.

7. The intelligent recognition and annotation method for endoscopic images as described in claim 1, characterized in that, Image annotation of the endoscopic standard image sequence using the continuous correction results includes: A graphic report is generated based on the image annotation results, and the graphic report is sent to the display terminal for visual interactive annotation display.

8. The intelligent recognition and annotation method for endoscopic images as described in claim 1, characterized in that, Reading the image sequence acquired during continuous advancement of the endoscope also includes: Perform imaging state verification detection on the image sequence and establish the verification detection results; If the verification result is a failure, a re-acquisition instruction is generated, and the image sequence is re-acquisitioned according to the re-acquisition instruction.

9. An intelligent recognition and annotation system for endoscopic images, characterized in that, The system is used to implement the intelligent recognition and annotation method for endoscopic images according to any one of claims 1-8, the system comprising: Image standardization module: Reads the image sequence acquired during the continuous advancement of the endoscope, performs brightness normalization and specular reflection suppression processing on each frame of the image sequence, and generates a standard image sequence of the endoscope with consistent imaging conditions; Feature extraction module: Based on the cavity edge contour, mucosal fold direction, and tissue texture periodic features, extract a set of structural features to characterize the distribution of anatomical structures within the endoscopic field of view, and generate an anatomical structure constraint map based on the set of structural features; Consistency Analysis Module: Within the spatial range defined by the anatomical structure constraint map, perform local directional texture consistency analysis on the standard endoscopic image sequence to establish abnormal texture perturbation components; Continuity analysis module: performs multi-scale aggregation on the abnormal texture perturbation components, establishes significant candidate images of lesions, and performs spatial position continuity analysis of adjacent frames of significant candidate images of lesions based on endoscopic standard image sequences, calculates position offset and morphological change, and constructs continuous correction results; Image annotation module: Performs image annotation of the standard endoscopic image sequence using the continuous correction results.

10. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable instructions; The processor, when executing executable instructions stored in the memory, implements the intelligent recognition and annotation method for endoscopic images as described in any one of claims 1-8.