Intelligent recognition and lesion positioning system for digestive endoscopic images
By constructing a lens pose change sequence and a texture displacement compensation mechanism, the problem of field of view variation in lesion localization in traditional gastrointestinal endoscopy was solved, achieving precise lesion localization and improving the accuracy of lesion contour tracking and spatial coordinate calculation.
Patent Information
- Application Number
- CN202610499624.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-15
- Publication Date
- 2026-07-10
AI Technical Summary
Traditional intelligent image recognition and lesion localization systems for digestive endoscopy struggle to establish a dynamic correlation mechanism between the physical movement of the lens and the viewing angle when the field of view changes drastically due to frequent turning and bending of the endoscope. This results in the loss of a fixed reference for the coordinate mapping of the lesion area in different image frames, causing tomographic tracking and localization deviations in the lesion contour.
The endoscope posture analysis module analyzes the lens posture change sequence, the image posture matching module filters the frame order and posture label pairing relationship, the mucosal texture displacement module compares the texture point position changes, the image rotation correction module compensates for the viewing angle offset, and the lesion coordinate positioning module adjusts the contour center relationship to achieve accurate positioning of the lesion spatial coordinates.
It effectively overcomes the defects of field of view changes and tissue spatial misalignment caused by the movement of the endoscope, improves the accuracy of dynamic contour tracking and spatial coordinate calculation for lesion localization, and ensures the baseline stability of boundary mapping in continuous images.
Smart Images

Figure CN122368183A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of endoscopic image analysis technology, and in particular to a system for intelligent recognition and lesion localization of digestive endoscopy images. Background Technology
[0002] The field of endoscopic image analysis technology mainly involves technologies that utilize imaging equipment to acquire images of the human digestive tract and to identify and analyze the content of these images. This typically includes endoscopic acquisition systems, image acquisition and storage, image enhancement processing, tissue structure recognition, lesion area identification, and medical image annotation. This field is widely used in gastroscopy, colonoscopy, and other digestive tract endoscopy examinations. Among these, the traditional intelligent recognition and lesion localization system for digestive endoscopic images refers to an image processing system that identifies lesions in images acquired during digestive endoscopic examinations and marks their locations within the images.
[0003] Traditional lesion identification mechanisms directly extract features and calculate positions based on a single static image during operation. This operating mode is easily affected by the drastic changes in the field of view caused by frequent turning and bending of the scope in actual operation. It is difficult to establish a dynamic correlation mechanism between the physical movement of the lens and the viewing angle of the image. This results in spatial misalignment of tissues and mucous membranes between consecutive images, causing the coordinate mapping of the lesion area in different image frames to lose its inherent reference support, thus leading to problems such as tomography in lesion contour tracking and increased positioning deviation. Summary of the Invention
[0004] To address the technical problems existing in the prior art, embodiments of the present invention provide an intelligent recognition and lesion localization system for digestive endoscopy images. The technical solution is as follows:
[0005] On the one hand, it provides an intelligent image recognition and lesion localization system for digestive endoscopy, including: The endoscope posture analysis module is based on electronic endoscope. It analyzes the sampling time corresponding to the voltage record of the angle detector, compares the continuous sampling changes, sorts out the correspondence between vertical bending and horizontal bending, and obtains the lens posture change sequence. The image pose matching module analyzes the video frame acquisition time and pose recording time based on the lens pose change sequence, filters adjacent pose records, adjusts the frame order and pose identifier pairing relationship, and obtains frame pose association identifier. The mucosal texture displacement module analyzes the changes in mucosal texture brightness arrangement in adjacent image frames based on the frame pose association identifier, compares the positions of the same texture point before and after, filters continuous texture points, and obtains the mucosal texture movement vector. The image rotation correction module compares the texture movement direction with the lens turning direction based on the mucosal texture movement vector, filters the rotation reference direction, and combines the cavity wall boundary with the image center position to obtain the viewing angle offset compensation amount. Based on the aforementioned viewpoint offset compensation amount, the lesion coordinate localization module analyzes the image chromaticity transition region, compares the brightness relationship between the inside and outside of the contour, filters continuous boundary pixels, adjusts the contour center relationship of adjacent image frames, and obtains the spatial contour coordinate data of the lesion.
[0006] On the other hand, the lens attitude change sequence includes lens pitch angle information, lens yaw angle information, and lens attitude time identifier; the frame attitude association identifier includes image frame number identifier, attitude time mapping identifier, and attitude matching index identifier; the mucosal texture movement vector includes texture point lateral displacement component, texture point longitudinal displacement component, and texture point movement direction identifier; the viewpoint offset compensation amount includes image rotation compensation angle, cavity wall center offset coordinates, and image center reference coordinates; and the lesion spatial contour coordinate data includes lesion contour center coordinates, lesion boundary pixel coordinates, and lesion region spatial range.
[0007] On the other hand, the endoscope posture analysis module includes: The voltage timing recognition submodule is based on an electronic endoscope. It analyzes the sampling time corresponding to the voltage record of the angle detector of the endoscope body turning control part, compares the time sequence relationship of adjacent sampling segments, determines the continuous state of the time interval of the sampling segments, identifies the continuous time sequence segments, and obtains the continuous voltage timing segments. The bending direction determination submodule obtains voltage change direction records based on the voltage time sequence continuous segments, compares the correspondence between the displacement change direction of the turning steel wire and the voltage change direction, determines the records with consistent directions in the continuous segments, filters the set of segments with consistent directions, and obtains bending direction sequence data. The posture sequence construction submodule obtains the posture change information corresponding to the direction record based on the bending direction sequence data, compares the time position relationship of the upper and lower bending records, determines the time corresponding state of the left and right bending records, adjusts the time mapping relationship between the upper and lower bending records and the left and right bending records, and obtains the lens posture change sequence.
[0008] On the other hand, the image pose matching module includes: The timing correspondence submodule, based on the lens attitude change sequence, calls the lens video frame acquisition timing sequence and attitude recording timing, compares the sequential relationship of the timing in the same order, verifies the corresponding spacing of adjacent timings, and adjusts the positions of the alignable timings to obtain the frame attitude timing mapping identifier. The neighbor filtering submodule, based on the frame attitude time mapping identifier, calls the adjacent time interval records before and after, compares the relationship between the frame time and the attitude record spacing, determines the time adjacency order, merges the spacing continuous records, and obtains the attitude adjacency matching identifier. The frame order matching submodule calculates the attitude position corresponding to each frame based on the attitude adjacency matching identifier, combined with the image frame number and the arrangement order of attitude records, and corrects the corresponding order of frame order and attitude identifier to obtain the frame attitude association identifier.
[0009] On the other hand, the mucosal texture displacement module includes: The brightness sorting submodule, based on the frame attitude association identifier, calls the pixel position corresponding to the mucosal region of adjacent image frames, compares the brightness order of pixels at the same position, verifies the correspondence between the brightness transfer direction and the time sequence between frames, and consolidates the brightness sorting record of the region to obtain the pixel brightness order value. The texture coherence submodule, based on the pixel brightness sequence, calls the previous and subsequent position records of the same texture point, compares the overlap relationship between the forward displacement trajectory and the back-pointing position trajectory, determines the cross-frame movement coherence status of the texture point, filters texture points with consistent position connections, and obtains the texture trajectory coherence rate. The displacement synthesis submodule calculates the trajectory direction corresponding to the texture point combination direction based on the texture trajectory continuity rate, combines the lateral displacement record and the longitudinal displacement record, determines the belonging relationship of the turning position of the jump trajectory, adjusts the classification of abnormal turning trajectory, and obtains the mucosal texture movement vector.
[0010] On the other hand, the image rotation correction module includes: The same direction discrimination submodule obtains the texture movement direction and lens turning record based on the mucosal texture movement vector, compares the arrangement order of the direction marks in each frame, verifies the consistent pointing relationship in the same frame, filters out the corresponding reverse records, and obtains the rotation direction reference value. Based on the rotation direction reference, the boundary positioning submodule calls the image edge brightness transition trajectory, compares the connection order of adjacent transition points, determines the cavity wall boundary closure direction, and organizes the continuous boundary position records to obtain the cavity wall boundary positioning amount. The center correction submodule, based on the cavity wall boundary positioning amount, checks the relative position of the cavity wall center and the image center, compares the correspondence between the center offset direction and the rotation direction, adjusts the center correspondence order, and obtains the viewpoint offset compensation amount.
[0011] On the other hand, the lesion coordinate localization module includes: Based on the aforementioned viewpoint offset compensation amount, the transition distribution submodule calls the colorimetric record of the digestive tract mucosal wall structure image, compares the colorimetric change order of adjacent pixels with the brightness transition relationship between the inner and outer sides of the contour, and adjusts the position of the continuous transition area to obtain the colorimetric transition distribution rate. Based on the chromaticity transition distribution rate, the boundary path submodule calls the boundary pixel adjacency record, compares the connection order of adjacent boundary pixels with the consistency of path turning, filters out continuous closed path segments, and straightens the connection direction to obtain the boundary closure path rate. Based on the boundary closure path rate, the contour connection submodule checks the corresponding position after the boundary coordinates are rotated, compares the displacement relationship between the contour centers of adjacent image frames, and adjusts the connection relationship between the contour centers of adjacent image frames to obtain the spatial contour coordinate data of the lesion.
[0012] On the other hand, the sampling time refers to the time marker corresponding to when the angle detector collects the voltage signal, the vertical bending refers to the voltage sampling record when the lens is bent in the vertical direction, and the horizontal bending refers to the voltage sampling record when the lens is bent in the horizontal direction.
[0013] On the other hand, the video frame acquisition time refers to the time sequence formed by arranging the time markers corresponding to each image frame in the electronic endoscope video stream in chronological order when they are acquired, and the attitude recording time refers to the time marker corresponding to the lens attitude angle recording during the sampling process.
[0014] On the other hand, the pairing relationship refers to the correspondence established between image frame acquisition records and posture records according to the time sequence, and the preceding and following positions refer to the spatial correspondence between the pixel positions corresponding to the same texture point in two consecutive image frames.
[0015] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: By analyzing the voltage trend of the angle detector in the control unit to construct a sequence of lens attitude changes, the attitude record and the acquisition sequence are paired to form an association identifier. The lens turning relationship is evaluated by combining the tracking mucosal texture displacement vector. Based on this, the viewing angle offset compensation amount is calculated and the cavity wall center position is corrected. In the chromatic transition area, the contour connection relationship between adjacent frames is adjusted. This effectively overcomes the field of view changes and tissue space misalignment defects caused by lens movement, ensures the baseline stability of boundary mapping in continuous images, and improves the accuracy of dynamic contour tracking and spatial coordinate calculation. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the system of the present invention; Figure 2 This is a schematic diagram of the system framework of the present invention; Figure 3 This is a flowchart of the endoscope posture analysis module of the present invention; Figure 4 This is a flowchart of the image pose matching module of the present invention; Figure 5 This is a flowchart of the mucosal texture displacement module of the present invention; Figure 6 This is a flowchart of the image rotation correction module of the present invention; Figure 7 This is a flowchart of the lesion coordinate localization module of the present invention. Detailed Implementation
[0018] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0019] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0020] This invention provides an intelligent image recognition and lesion localization system for digestive endoscopy, such as... Figure 1 As shown, the system includes: The endoscope attitude analysis module is based on electronic endoscope. It analyzes the corresponding sampling time of the voltage record of the angle detector of the endoscope body steering control unit, compares the continuity relationship of the sampling segments before and after, calculates the voltage change trend corresponding to the displacement of the steering wire, filters the time sequence connection segments, adjusts the correspondence between the up and down bending records and the left and right bending records, and obtains the lens attitude change sequence. The image pose matching module analyzes the correspondence between the acquisition time sequence of the lens video frame and the pose recording time based on the lens pose change sequence, compares the state of the interval between adjacent time moments, filters the neighboring pose records, calculates the pose position corresponding to the image frame number, adjusts the pairing relationship between the frame order and the pose identifier, and obtains the frame pose association identifier. The mucosal texture displacement module analyzes the changes in pixel brightness arrangement in the mucosal region of adjacent image frames based on frame pose association identifiers, compares the corresponding relationship of the same texture point before and after, filters texture points with continuous trajectories, calculates the combined direction of lateral and vertical displacements, adjusts the affiliation relationship of jump trajectories, and obtains the mucosal texture movement vector. The image rotation correction module is based on the mucosal texture movement vector to obtain the texture movement direction and lens rotation record, compares whether the texture movement direction and lens rotation are in the same direction, filters the available rotation reference direction, calculates the cavity wall boundary position corresponding to the brightness transition trajectory of the image edge, adjusts the relative relationship between the cavity wall center and the image center, and obtains the viewpoint offset compensation amount. The lesion coordinate localization module analyzes the distribution of image color transition regions of the digestive tract mucosal wall structure based on the viewpoint offset compensation amount, compares the brightness change relationship between the inner and outer regions of the contour, filters the continuous boundary pixel connection path, calculates the corresponding position after the boundary coordinate rotation, and adjusts the connection relationship of the contour center of adjacent image frames to obtain the spatial contour coordinate data of the lesion.
[0021] The lens attitude change sequence includes lens pitch angle information, lens yaw angle information, and lens attitude time marker. The frame attitude association marker includes image frame number marker, attitude time mapping marker, and attitude matching index marker. The mucosal texture movement vector includes texture point lateral displacement component, texture point longitudinal displacement component, and texture point movement direction marker. The viewpoint offset compensation amount includes image rotation compensation angle, cavity wall center offset coordinates, and image center reference coordinates. The lesion spatial contour coordinate data includes lesion contour center coordinates, lesion boundary pixel coordinates, and lesion region spatial range.
[0022] In the endoscope attitude analysis module, the endoscope body turning control unit refers to the mechanical control structure on the electronic endoscope handle used to control the bending direction of the lens. It drives the internal traction structure of the endoscope to generate displacement through a knob or lever, thereby changing the pointing of the lens end. The sampling moment refers to the time mark corresponding to when the angle detector collects the voltage signal. This time mark is used to indicate the time sequence of voltage record generation. The continuous relationship of sampling segments refers to the voltage records between adjacent sampling moments being arranged continuously on the time axis, used to reflect the signal changes during the same lens turning action. The turning wire refers to the traction structure set inside the endoscope and connected to the endoscope body turning control unit. This structure is displaced when the control unit is activated, causing the lens end to bend. The voltage change trend refers to the change trend of the voltage signal output by the angle detector in the time series. This change trend is used to indicate the change in the lens bending direction. The time sequence connection segment refers to a group of continuous voltage record segments arranged in chronological order without sampling interruption, used to describe a continuous lens attitude change process. The vertical bending record refers to the voltage sampling record used to represent the lens bending state in the vertical direction. The horizontal bending record refers to the voltage sampling record used to represent the lens bending state in the horizontal direction.
[0023] In the image pose matching module, the video frame acquisition time sequence refers to the time sequence formed by arranging the time markers corresponding to the acquisition of each image frame in the electronic endoscope video stream in chronological order; the pose recording time refers to the time marker corresponding to the lens pose angle recording during the sampling process; the interval state refers to the spacing relationship between two adjacent time markers on the time axis, used to determine the order and proximity between time records; the adjacent pose record refers to the pose record that is close to the acquisition time of a certain image frame on the time axis; the pose position refers to the position index corresponding to a certain image frame in the pose recording sequence; and the pairing relationship refers to the one-to-one correspondence established between the image frame acquisition record and the pose record based on the time order.
[0024] In the mucosal texture displacement module, the mucosal region refers to the image area in the endoscopic image that presents the tissue structure of the digestive tract wall, which includes mucosal texture structure and vascular texture structure; the front-back position correspondence refers to the spatial correspondence between the pixel positions of the same texture point in two consecutive frames of images; the trajectory continuous texture point refers to the texture feature point that can be continuously identified in multiple frames of images and maintains the continuity of spatial position change; the combined direction refers to the two-dimensional movement direction formed by the lateral and longitudinal position changes of the texture point; the jump trajectory attribution relationship refers to the classification relationship of removing trajectories that do not conform to the continuous change law from the texture point set after judging the texture point displacement trajectory.
[0025] In the image rotation correction module, the available rotation reference direction refers to the reference direction selected for image rotation correction after comparing the lens turning direction with the texture movement direction; image edge brightness refers to the spatial distribution of pixel brightness in the image boundary area, which is used to identify the edge position of the inner wall of the cavity; cavity wall boundary position refers to the edge contour position formed by the inner wall of the digestive tract cavity in the endoscopic image; relative relationship refers to the positional difference between the cavity wall center position and the image center position in spatial coordinates.
[0026] In the lesion coordinate localization module, the image color transition region distribution refers to the spatial distribution state of the pixel color or brightness changes and forms a transition region in the endoscopic image; the brightness change relationship refers to the difference in pixel brightness between the lesion area and the surrounding mucosal area; the continuous boundary pixel connection path refers to the path structure that connects adjacent boundary pixels according to spatial adjacency to form a closed contour; the boundary coordinate rotation refers to the spatial rotation transformation processing of the lesion contour pixel coordinates based on rotation compensation information; and the connection relationship refers to the continuous change relationship of the lesion contour center position in spatial coordinates in adjacent image frames.
[0027] like Figure 2 and Figure 3 As shown, the endoscope posture analysis module includes: The voltage timing recognition submodule is based on an electronic endoscope. It analyzes the sampling time corresponding to the voltage record of the angle detector of the endoscope body turning control part, compares the time sequence relationship of adjacent sampling segments, determines the continuous state of the time interval of the sampling segments, identifies the continuous time sequence segments, and obtains the continuous voltage timing segments. The voltage recording time stamp processing based on the internal angle detector of the electronic endoscope control handle first reads a set of continuous sampling records, for example, in the time sequence of 0.000 seconds, 0.020 seconds, 0.041 seconds, 0.061 seconds, and 0.083 seconds. The time difference between adjacent sampling records is calculated sequentially, yielding four values: 0.020 seconds, 0.021 seconds, 0.020 seconds, and 0.022 seconds. Then, a preset time continuity threshold is set to 0.030 seconds, based on the common 20-millisecond sampling period of the electronic endoscope control circuit with an allowance of 10-millisecond fluctuation. Subsequently, the time difference is interval-wise judged. When the time difference value is between 0 seconds and 0.030 seconds, it is marked as continuous sampling; when the time difference is greater than 0.030 seconds, it is marked as discontinuous. For example, a time difference of 0.085 seconds is classified as discontinuous. Then, the marked states are read sequentially according to the time sequence. When the number of consecutive time differences... When there are three or more records, they are grouped into the same continuous segment. For example, if four records between 0.000 seconds and 0.083 seconds all meet the continuity condition, then this segment is written into segment set A. If a subsequent record has a time difference of 0.085 seconds, then the recording of set A is terminated and a new set B is started. During the set recording process, the start time and end time of the segment are recorded simultaneously. For example, the start time of set A is 0.000 seconds and the end time is 0.061 seconds. Then, the duration is calculated by subtracting the start time from the end time to get 0.061 seconds. The duration is then compared with the set minimum effective segment threshold of 0.040 seconds. If the duration is greater than or equal to 0.040 seconds, the segment record is retained. If the duration is less than 0.040 seconds, the segment is discarded. For example, if a segment has a duration of only 0.019 seconds, it is directly deleted. Finally, the retained continuous segment set is written into the recording table in chronological order to form voltage timing continuous segments.
[0028] The bending direction determination submodule obtains voltage change direction records based on continuous voltage time segments, compares the correspondence between the displacement change direction of the turning wire and the voltage change direction, determines the records with consistent directions in the continuous segments, filters the set of segments with consistent directions, and obtains bending direction sequence data. Extract voltage values from the same continuous segment. For example, read voltage records of 2.45V, 2.53V, 2.61V, 2.58V, and 2.55V. Then, perform difference calculations on adjacent records: the first group is 2.53 minus 2.45, resulting in 0.08V; the second group is 2.61 minus 2.53, resulting in 0.08V; the third group is 2.58 minus 2.61, resulting in -0.03V; and the fourth group is 2.55 minus 2.58, resulting in -0.03V. Then, classify the records into two categories based on the sign of the difference: positive and negative changes. Differences greater than 0 are classified as positive changes, and differences less than 0 are classified as negative changes. Simultaneously, read the corresponding steering wire displacement records for that time period, for example, displacement values of 0.32mm, 0.40mm, 0.48mm, 0.46mm, and 0.43mm. Then, perform position... The displacement values are calculated to be 0.08 mm, 0.08 mm, -0.02 mm, and -0.03 mm. Then, the signs of the voltage change direction and displacement change direction are checked one by one. If the signs are the same, it is marked as a record with consistent direction. If the signs are different, it is marked as a record with inconsistent direction. For example, if 0.08 volts and 0.08 mm are both positive changes, it is judged as a consistent record. If the signs are different, it is marked as an inconsistent record. Then, the number of consecutive consistent records is counted and the threshold for consistent direction is set to 3. When the number of consecutive consistent records reaches 3 or more, the time segment is written into the set of consistent direction segments. For example, if the first 3 sets of records are all consistent changes, they are written into set A. If the number of consecutive consistent records is less than 3, they are not written into the set. After completing the traversal of all records, the sets are arranged in chronological order to form the bending direction sequence data.
[0029] The attitude sequence construction submodule obtains the attitude change information corresponding to the direction record based on the bending direction sequence data, compares the time position relationship of the upper and lower bending records, determines the time corresponding state of the left and right bending records, adjusts the time mapping relationship between the upper and lower bending records and the left and right bending records, and obtains the lens attitude change sequence. Extract voltage sampling records for vertical bends, for example, time sequences of 0.10 seconds, 0.12 seconds, and 0.14 seconds, with voltage values of 2.30 V, 2.36 V, and 2.41 V respectively. Simultaneously extract voltage sampling records for horizontal bends, for example, time sequences of 0.11 seconds, 0.13 seconds, and 0.15 seconds, with voltage values of 1.92 V, 1.98 V, and 2.05 V respectively. Then, calculate the time difference between the vertical and horizontal records, for example, the time difference between 0.12 seconds and 0.11 seconds is 0.01 seconds, and the time difference between 0.14 seconds and 0.13 seconds is also 0.01 seconds. Set the time matching threshold to 0.020 seconds. When the time difference is less than or equal to 0.020 seconds, mark it as a time-corresponding record; when the time difference is greater than 0.020 seconds, mark it as a non-corresponding record. After obtaining the corresponding records, calculate the voltage sampling records for vertical bends separately. The voltage changes during bending and lateral bending are recorded. For example, the vertical voltage changes from 2.30 volts to 2.41 volts by 0.11 volts, and the horizontal voltage changes from 1.92 volts to 2.05 volts by 0.13 volts. Then, the proportion of these two sets of changes in the overall change is calculated. The calculation process is as follows: add 0.11 and 0.13 to get 0.24, then divide 0.11 by 0.24 to get 0.458, and divide 0.13 by 0.24 to get 0.542. Record 0.458 in the vertical proportion record and 0.542 in the horizontal proportion record, while also recording the corresponding time markers. For example, 0.12 seconds corresponds to a vertical proportion of 0.458 and a horizontal proportion of 0.542. Then, the next set of records is read and the above steps are repeated. Finally, all records are arranged in chronological order to form a sequence of lens posture changes.
[0030] like Figure 2 and Figure 4 As shown, the image pose matching module includes: The timing correspondence submodule is based on the lens attitude change sequence. It calls the lens video frame acquisition timing sequence and attitude recording timing, compares the sequential relationship of the timing in the same order, verifies the corresponding spacing of adjacent timings, and adjusts the positions of the alignable timings to obtain the frame attitude timing mapping identifier. Read consecutive time points from the attitude recording table, such as 0.100 seconds, 0.133 seconds, 0.166 seconds, and 0.199 seconds. Then read the corresponding time points from the video frame acquisition table, such as 0.102 seconds, 0.136 seconds, 0.168 seconds, and 0.201 seconds. Merge and sort the two sets of times in chronological order, then check for reverse order mapping points. If the previous frame time is earlier than the next frame time and the corresponding attitude times maintain the same order, write a consistent order marker at that position. If frame times or attitude times are reversed, mark that position as an abnormal position and stop the alignment for that group. Then continue reading the time difference between adjacent items, for example, video frame intervals of 0.034 seconds, 0.032 seconds, and 0.033 seconds, and attitude recording intervals of 0.033 seconds, 0.033 seconds, and 0.033 seconds. Then merge the two sets of times in chronological order. The interval values are calculated item by item to obtain 0.001 seconds, 0.001 seconds, and 0 seconds, and compared with the interval verification threshold of 0.008 seconds. This threshold is set according to the single frame period of approximately 0.033 seconds corresponding to 30 frames per second of endoscopy video, allowing for a deviation of about one-quarter of a frame. When the difference is in the range of 0 seconds to 0.008 seconds, the position is retained; when the difference is greater than 0.008 seconds, the position is discarded. For example, if a group of frames has an interval of 0.050 seconds and an attitude interval of 0.033 seconds, the difference is 0.017 seconds, which is directly marked as an unaligned record. Then, the continuously retained positions are organized by sequence number to form a mapping queue in which the frame sequence number corresponds one-to-one with the attitude time sequence number. For example, the 15th frame corresponds to the 12th attitude record, the 16th frame corresponds to the 13th attitude record, and the 17th frame corresponds to the 14th attitude record. The frame attitude time mapping identifier after being organized by time position is output.
[0031] The neighbor filtering submodule is based on the frame attitude time mapping identifier, calls the adjacent time interval records, compares the relationship between the frame time and the attitude record spacing, determines the time adjacency order, merges the consecutive spacing records, and obtains the attitude adjacency matching identifier. Read the frame time and pose time difference values before and after each mapped position. For example, if the current frame time is 0.500 seconds and the corresponding pose time is 0.497 seconds, the time difference between the previous frame and the previous pose is 0.004 seconds, and the time difference between the next frame and the next pose is 0.006 seconds. Then, compare the current difference, the forward difference, and the backward difference item by item. First, determine whether the current difference falls within the adjacency allowable range of 0 seconds to 0.010 seconds. This threshold is set according to the common small drift between the video frame sampling period and the pose sampling period. If the current difference exceeds 0.010 seconds, remove the record from the adjacency judgment sequence. If the current difference does not exceed 0.010 seconds, continue to check whether the order of the previous and next positions is continuous, that is, whether the previous position number, the current position number, and the next position number are each incremented by 1. For example, the 20th frame corresponds to the 18th pose. If frame 21 corresponds to pose 19 and frame 22 corresponds to pose 20, then the adjacency order is considered valid. If frame 21 directly corresponds to pose 21, then poses 19 and 20 are missing in between, which is recorded as an order jump. Subsequently, records that meet the adjacency conditions are merged. Records with continuous time difference fluctuations not exceeding 0.003 seconds are written into the same adjacency group. For example, 0.004 seconds, 0.005 seconds, and 0.006 seconds are grouped together, while 0.004 seconds, 0.011 seconds, and 0.005 seconds are cut off at 0.011 seconds. During merging, the start and end positions of the record group are synchronized. For example, frames 20 to 28 form a group, and frame 29 is removed due to a sudden increase in the difference. After completing all position checks, the remaining adjacent groups are numbered sequentially to generate pose adjacency matching identifiers.
[0032] The frame order matching submodule calculates the attitude position corresponding to each frame based on the attitude adjacency matching identifier, combined with the image frame number and the arrangement order of attitude records, and corrects the corresponding order of frame order and attitude identifier to obtain the frame attitude association identifier. Read the sequence numbers of each image frame within the same adjacent group, such as frames 120, 121, 122, and 123, and simultaneously read the corresponding attitude record sequence numbers within the same group, such as records 85, 86, 87, and 88. Then, calculate the increment of the frame sequence number and the increment of the attitude sequence number for each item. The former will be 1, 1, 1, and the latter will also be 1, 1, 1. Compare the two sets of increments one by one. If the difference between the positions in the same sequence is 0, it is marked as a consistent pairing order. If a position shows a frame sequence number increment of 1 and an attitude sequence number increment of 2, it is recorded as an attitude jump. If a position shows a frame sequence number increment of 2 and an attitude sequence number increment of 1, it is recorded as a frame sequence gap. Then, perform correction on the abnormal positions. During correction, first read the valid pairing before and after the abnormal point. For example, frame 121 corresponds to attitude record 86. Frame 123 corresponds to pose 88, while frame 122 is missing a pair. Therefore, frame 122 is directly assigned to pose 87. If an abnormal segment is missing 2 or fewer consecutive frames, interpolation is allowed; if more than 2 frames are missing, interpolation stops. Then, each frame and its corresponding pose position are written into the pairing table. The pairing table records the frame number, pose number, time difference, and adjacency group number. For example, frame 120 corresponds to pose 85 with a time difference of 0.003 seconds and is in the 3rd adjacency group. Frame 121 corresponds to pose 86 with a time difference of 0.004 seconds and is also in the 3rd adjacency group. After all adjacency groups have been corrected, the entire table is checked again to see if the frame number and pose number are continuously increasing and if the adjacent time difference is within the range of 0 seconds to 0.010 seconds. After completion, the frame pose association identifier is output.
[0033] like Figure 2 and Figure 5 As shown, the mucosal texture displacement module includes: The brightness sorting submodule, based on the frame attitude association identifier, calls the pixel position corresponding to the mucous membrane region of adjacent image frames, compares the brightness order of pixels at the same position, verifies the correspondence between the brightness transfer direction and the time sequence between frames, and consolidates the brightness sorting record of the region to obtain the pixel brightness order value. In continuous endoscopic video images, fixed pixel positions in the mucosal region are selected for brightness reading. For example, in frame 180, the grayscale brightness values at coordinates (320, 240), (321, 240), and (322, 240) are read as 138, 142, and 147, respectively. Simultaneously, in the next frame, frame 181, the brightness values at the same pixel positions are read as 141, 145, and 149. Then, the brightness difference between the preceding and following frames is calculated sequentially, yielding 141 - 138 = 3, 145 - 142 = 3, and 149 - 147 = 2. These differences are recorded as brightness transfer values and written to a change record table. Subsequently, the corresponding frame time records are read; for example, frame 180 has a time of 6.000 seconds, and frame 181 has a time of 6.033 seconds. The direction of brightness change is then compared item by item with the frame time sequence. When the brightness value of a later frame is greater than that of a previous frame and the time sequence increases, the change is considered complete. The brightness changes are recorded sequentially. When the brightness decreases and the time increases, it is recorded as a reverse brightness change. Then, the brightness value of the 182nd frame is read, for example, 143, 147, 151. The difference is calculated again to get 2, 2, 2. The brightness values of the same pixel position in the three frames are arranged according to the value. For example, 138, 141, 143 form an increasing sequence, 142, 145, 147 form an increasing sequence, and 147, 149, 151 form an increasing sequence. The brightness stability judgment interval is set to three consecutive frames with a difference between 1 and 6 and the same direction of change. When the condition of this interval is met, the pixel is recorded as a valid brightness sequence. Then, the reading, difference calculation, sequential judgment and arrangement record are repeated for dozens of pixel positions in the same area. All valid brightness sequences are labeled with the sequence number according to the frame order and organized into a regional brightness arrangement record. The pixel brightness sequence value is output.
[0034] The texture coherence submodule, based on the pixel brightness sequence, calls the previous and subsequent position records of the same texture point, compares the overlap relationship between the forward displacement trajectory and the back-pointing position trajectory, determines the cross-frame movement coherence of the texture point, filters texture points with consistent position connection, and obtains the texture trajectory coherence rate. The system reads the coordinate changes of the same texture point in consecutive image frames. For example, the position is (415, 266) in frame 240, (417, 267) in frame 241, and (419, 268) in frame 242. First, it calculates the forward displacement difference. The horizontal difference between frame 241 and frame 240 is 417 minus 415, which equals 2, and the vertical difference is 267 minus 266, which equals 1. Simultaneously, it calculates the horizontal difference between frame 242 and frame 241: 419 minus 417, which equals 2, and the vertical difference is 268 minus 267, which equals 1. Then, it performs a back-pointing check, subtracting the displacement difference from the previous frame from the coordinates of frame 242 to obtain 417 and 267. These are then compared item by item with the coordinates of frame 241. When the two coordinate values are completely identical, it is recorded as a trajectory overlap record. Subsequently, the same back-pointing check is performed on frame 241. The operation involves subtracting 2 from 417 to get 415, and subtracting 1 from 267 to get 266. These are then compared with the coordinates of the 240th frame. If the coordinates match, they are recorded as overlapping records. The number of times the current texture point overlaps in three consecutive frames is then counted. For example, if there are two overlaps, the number of overlaps is 2. This value is then compared with the total number of consecutive checks, 2, to get 1.0. This value is then compared with the set coherence threshold of 0.75. If the value is between 0.75 and 1.00, it is marked as a coherent texture point. If the value is less than 0.75, it is marked as a broken texture point. Subsequently, the same position reading, displacement difference calculation, and back-pointing check operations are performed on all texture points in the entire mucosal region. For example, if 96 out of 120 texture points meet the coherence condition, the coherence ratio of the recorded region is 0.80, and the texture trajectory coherence rate is output.
[0035] The displacement synthesis submodule calculates the trajectory direction corresponding to the combination of texture points based on the texture trajectory coherence rate, combines the lateral displacement record and the longitudinal displacement record, determines the belonging relationship of the turning position of the jump trajectory, adjusts the classification of abnormal turning trajectory, and obtains the mucosal texture movement vector. The horizontal and vertical displacement changes of the same texture point are read in consecutive image frames. For example, from frame 310 to frame 311, the horizontal displacement is 3 pixels and the vertical displacement is 2 pixels; from frame 311 to frame 312, the horizontal displacement is 3 pixels and the vertical displacement is 1 pixel. Then, a combined displacement calculation is performed on each set of horizontal and vertical displacements. This is done by first calculating the square of the horizontal displacement value, then the square of the vertical displacement value, summing the two, and then taking the square root to obtain the displacement amplitude. For example, 3 squared is 9, 2 squared is 4, adding them together gives 13, and then taking the square root gives approximately 3.61 pixels. Simultaneously, the trajectory direction is determined based on the signs of the horizontal and vertical displacements. When both the horizontal and vertical displacements are positive, it is recorded as the lower right direction; when both are positive, it is recorded as the upper right direction; when both are negative, it is recorded as the lower left direction; and when both are negative, it is recorded as the lower left direction. Negative values are recorded as the top-left direction. For example, a horizontal change of 3 and a vertical change of 2 are recorded as the bottom-right direction. Then, the changes in the next frame are read, and the changes of 3 and 1 are recorded again as the bottom-right direction. The number of consecutive direction records is counted. For example, if two consecutive direction records are the same, the direction count is 2. This is compared with a set stability threshold of 2. When the number of direction records is greater than or equal to 2, the trajectory is marked as a stable movement trajectory. Then, the changes in the next frame are read. For example, if the horizontal change is negative 2 and the vertical change is 3, the direction changes to the bottom-left direction. This position is recorded as the trajectory turning point. The difference in displacement amplitude before and after the turning point is compared. For example, the difference between 3.61 pixels and 3.60 pixels is 0.01 pixels. When the difference is in the range of 0 to 0.50 pixels, it is classified into the same trajectory. When the difference is greater than 0.50 pixels, it is classified into an abnormal trajectory. Then, the direction judgment and turning point verification processing is repeated for all texture points, and the mucosal texture movement vector is output.
[0036] like Figure 2 and Figure 6 As shown, the image rotation correction module includes: The same direction discrimination submodule is based on the mucosal texture movement vector to obtain the texture movement direction and lens turning record, compares the arrangement order of the direction markers in each frame, verifies the consistent pointing relationship in the same frame, filters out the corresponding reverse records, and obtains the rotation direction reference value. The process involves acquiring records of texture movement direction in consecutive image frames. For example, the direction is lower right in frame 150, lower right in frame 151, and lower right in frame 152. Simultaneously, it reads the camera rotation direction recorded by the camera control unit, for example, lower right in frame 150, lower right in frame 151, and upper right in frame 152. Then, a direction consistency check is performed frame by frame, comparing the texture direction with the camera rotation direction. When the two direction markers are completely identical, a consistency record is written at the corresponding frame position; when the directions are different, a reverse record is written. For example, frames 150 and 151 are both recorded as consistent, while frame 152 is recorded as reversed. The number of consistent records in consecutive frames is then counted. For example, if 4 out of 5 consecutive frames are consistent, the consistency ratio is calculated as 4 divided by 5, resulting in 0.80. This value is then compared with the set same-direction threshold of 0.70. When the ratio is between 0.70 and 1.00, the segment of the record is included in the valid same-direction record. When the ratio is less than 0.70, the segment is removed from the reference sequence. Subsequently, the direction reading, frame-by-frame verification, and ratio statistics operations are performed segment by segment in the entire video frame sequence. All frame records that meet the same-direction condition are reordered according to frame order to form a direction reference sequence. At the same time, frame records that do not meet the condition are marked as reverse corresponding records and deleted from the sequence. The remaining frame direction records are then organized to form a rotation direction reference.
[0037] The boundary localization submodule uses rotation direction reference to call the image edge brightness transition trajectory, compares the connection order of adjacent transition points, determines the cavity wall boundary closure direction, and organizes continuous boundary position records to obtain the cavity wall boundary localization amount. Extract edge brightness change points from the corresponding frame image. For example, extract edge transition pixels (120, 320), (122, 321), (124, 322), and (126, 323) from the 210th frame image. Simultaneously, read the grayscale value changes at these locations. For instance, on the same edge line, the pixel brightness changes from 85 to 120 and then to 160. Then, connect adjacent pixels according to their coordinates and record the connection order. For example, (120, 320) connects to (122, 321), (122, 321) connects to (124, 322), and (124, 322) connects to (126, 323). Subsequently, perform a closure direction check on the connected point sequence. First, calculate the horizontal and vertical differences between adjacent points. For example, subtract 120 from 122 to get... 2,321 minus 320 equals 1. The same calculation is then performed on the next group to obtain 2 and 1. When the difference remains stable and the direction is consistent, it is recorded as a continuous boundary segment. Then, new transition points are read, such as (128, 324) and (130, 325), and the same difference calculation is performed. When the connection sequence returns to the vicinity of the starting position, such as the coordinates being close to (120, 320) and the distance is less than 5 pixels, it is recorded as a closed boundary segment. Then, the number of boundary points in the closed segment is counted. For example, 40 consecutive boundary points are recorded to form a complete cavity wall boundary record. Then, all boundary points are written into the boundary position table according to the connection order, and discrete points that cannot form a continuous connection are deleted, such as edge pixels that only have single or double point records. The continuous boundary point sequence is sorted out and the cavity wall boundary positioning value is output.
[0038] The center correction submodule, based on the cavity wall boundary positioning amount, checks the relative position of the cavity wall center and the image center, compares the correspondence between the center offset direction and the rotation direction, adjusts the center correspondence order, and obtains the viewpoint offset compensation amount. Read the coordinates of all boundary points in the same image frame. For example, if the boundary point set consists of 40 consecutive points such as (120, 320), (122, 321), (124, 322), and (126, 323), then sum the x-coordinates of all boundary points. For example, the sum of the 40 x-coordinates is 4980, which is then divided by the number of boundary points (40) to get 124.5. Simultaneously, sum the coordinates of all y-coordinates. For example, the sum of the 40 y-coordinates is 12980, which is then divided by 40 to get 324.5. Record the resulting coordinates (124.5, 324.5) as the cavity wall center position. Then read the center coordinates of the current image frame. For example, if the image resolution is 640×480, the center position is (320, 240). Finally, calculate the lateral and longitudinal offsets between the cavity wall center and the image center. For example, the lateral offset is... Subtracting 124.5 from 320 yields 195.5. Subtracting 324.5 from the vertical offset yields -84.5. Then, the rotation direction recorded in the rotation direction reference is read. For example, if the current frame's rotation direction is the lower right, the offset direction is checked to see if it matches the rotation direction. When the horizontal offset is positive and the vertical offset is negative, it is recorded as an upper right offset. When both the horizontal and vertical offsets are positive, it is recorded as a lower right offset. For example, 195.5 and -84.5 correspond to an upper right offset, which is inconsistent with the current lower right rotation direction. Therefore, it is marked as a sequential offset record in the offset record table. Subsequently, the same center calculation and offset direction judgment are performed on adjacent frames, and the offset records are reordered according to the frame order. When the offset direction matches the rotation direction for three consecutive frames, the offset is recorded as valid correction data. All valid offset records are then compiled to form the view offset compensation amount.
[0039] like Figure 2 and Figure 7 As shown, the lesion coordinate localization module includes: The transition distribution submodule, based on the viewpoint offset compensation amount, calls the colorimetric record of the digestive tract mucosal wall structure image, compares the colorimetric change order of adjacent pixels with the brightness transition relationship between the inner and outer sides of the contour, and adjusts the position of the continuous transition area to obtain the colorimetric transition distribution rate. In the current endoscopic image, a continuous pixel band within the candidate lesion region is selected, and the chromaticity and luminance records are read point by point. For example, in the 260th frame image, seven pixels with coordinates (280, 210), (281, 210), (282, 210), (283, 210), (284, 210), (285, 210), and (286, 210) are read sequentially along the horizontal direction. Their chromaticity values are 96, 101, 109, 118, 126, 129, and 131, respectively. The corresponding brightness values are read as 142, 145, 149, 121, 116, 112, and 109. Then, the chromaticity difference between adjacent pixels is calculated one by one, yielding 5, 8, 9, 8, 3, and 2. These differences are then compared one by one with the permissible chromaticity change range of 1 to 12. Differences within this range are marked as valid change records, while differences greater than 12 or less than 1 are marked as abrupt change records. Subsequently, a brightness transition check is performed on the same pixel sequence. First, the brightness of the first three pixels is summed to 142 plus 14. Adding 149 to 5 gives 436, then dividing by 3 gives 145.3. Next, summing the brightness of the last three pixels (116 + 112 + 109) gives 337, then dividing by 3 gives 112.3. The difference between the mean values of the first and last three pixels (33.0) is compared with the brightness transition threshold of 18. When the difference is greater than 18, it is recorded as an inner / outer brightness transition. Subsequently, consecutive pixel segments with valid chromaticity transitions and valid brightness transitions are grouped into the same transition region. For example, (280, 210) to (286, 210) are recorded as the first... A transition segment is defined, and then the chroma reading, difference comparison, and luminance mean check operations are repeatedly performed at the vertical and diagonal adjacent positions in this area. If the number of consecutive pixels is less than 4, the segment record is deleted. If the number of consecutive pixels is 4 or more, it is retained and recorded as a valid transition area. Then, the number of all valid transition pixels in the candidate area is counted. For example, if the total number of pixels in the area is 900 and the number of valid transition pixels is 243, 243 is divided by 900 to get 0.27 and recorded as the chroma transition distribution rate of the current image frame.
[0040] The boundary path submodule, based on the chromaticity transition distribution rate, calls the boundary pixel adjacency record, compares the connection order of adjacent boundary pixels with the consistency of path turning, filters out continuous closed path segments, and straightens the connection direction to obtain the boundary closure path rate. Read the adjacent records of boundary pixels in the candidate lesion region and arrange them in spatial order. For example, read the consecutive boundary pixels (301, 220), (302, 221), (303, 222), (304, 223), (304, 224), (303, 225), and (302, 226). Then, check the connection relationship between adjacent pixels one by one. First, calculate the horizontal difference and vertical difference between adjacent pixels. For example, the horizontal difference between (301, 220) and (302, 221) is calculated. A horizontal difference of 1 and a vertical difference of 1 indicate a valid adjacent connection. Similarly, (304, 223) and (304, 224) with a horizontal difference of 0 and a vertical difference of 1 are also recorded as a valid connection. If either the horizontal or vertical difference is greater than 2, the connection is marked as broken, and the current path recording is terminated. Subsequently, a direction consistency check is performed on the already formed pixel connection sequence. First, the direction of the first segment is recorded as the lower right direction. Then, the direction of the next segment is read as either the lower right or downward direction. When the direction change does not exceed [a certain value], [further details are provided]. When a direction level is crossed, it is recorded as a continuous turn. When the direction suddenly changes from the lower right to the upper left, it is recorded as a reverse turn. Then, the number of reverse turns in the path segment is counted and compared with the turn threshold 2. If the number of turns does not exceed 2, the path segment is retained. If it exceeds 2, the path segment is split into two independent path records. Then, new boundary pixels are read along the adjacent direction. When the distance between the end pixel and the start pixel is not greater than 2 pixels, it is recorded as a closed path segment. For example, if the horizontal difference between the end pixel (300, 221) and the start pixel (301, 220) is 1 and the vertical difference is 1, it is recorded as a closed path segment. Then, the number of boundary pixels in the closed path segment is counted. For example, if there are 52 consecutive boundary pixels, 52 is compared with the closure threshold 30. If it is greater than 30, the path segment is retained. If it is less than 30, it is deleted. Finally, the number of closed path segments and the number of candidate path segments are counted. For example, if there are 8 candidate path segments and 6 closed path segments are retained, 6 is divided by 8 to get 0.75 and recorded as the boundary closure path rate.
[0041] The contour connection submodule, based on the boundary closure path rate, checks the corresponding positions after the boundary coordinates are rotated, compares the displacement relationship between the contour centers of adjacent image frames, and adjusts the connection relationship between the contour centers of adjacent image frames to obtain the spatial contour coordinate data of the lesion. Extract the set of closed boundary coordinates for each frame in a series of image frames and calculate the corresponding contour center position. For example, in frame 340, read 40 boundary points (318, 232), (334, 232), (334, 244), and (318, 244). Sum all the horizontal coordinates to get 4980 and divide by 40 to get 124.5. Sum all the vertical coordinates to get 12980 and divide by 40 to get 324.5. Thus, the contour center coordinates of frame 340 are (124.5, 324.5). Then, read the contour center coordinates of frame 341 after rotation compensation, for example, (127.5, 326.5). Then read the contour center coordinates of frame 342 as (129.5, 327.5). Then, calculate the displacement of adjacent contour centers frame by frame. The horizontal displacement of frame 341 relative to frame 340 is 3 and the vertical displacement is 2. The horizontal displacement of frame 342 relative to frame 341 is... 2. The vertical displacement is set to 1, and the displacement value is compared with the allowable range of continuous displacement. The allowable range of horizontal displacement is set to 0 to 6 pixels, and the allowable range of vertical displacement is set to 0 to 5 pixels. When the displacement is within this range, it is recorded as a valid connection record. If the horizontal displacement exceeds 8 pixels or the vertical displacement exceeds 7 pixels, it is marked as a jump record and the contour data of that frame is removed. Then, the corresponding positions of the boundary points of each frame after rotation are checked. For example, the horizontal difference of the boundary point (318, 232) of frame 340 and the corresponding point (321, 234) of frame 341 is 3 and the vertical difference is 2. If both are within the allowable range, they are recorded as valid corresponding points. If the deviation of a boundary point exceeds 8 pixels, the point is deleted and the contour center position is recalculated. When the contour center displacement of three consecutive frames is within the allowable range, the frame records are grouped into the same spatial contour segment. The contour center coordinates, boundary pixel coordinates and regional range data of each frame are organized to form the lesion spatial contour coordinate data.
[0042] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A system for intelligent recognition and lesion localization of digestive endoscopy images, characterized in that, The system includes: The endoscope posture analysis module is based on electronic endoscope. It analyzes the sampling time corresponding to the voltage record of the angle detector, compares the continuous sampling changes, sorts out the correspondence between vertical bending and horizontal bending, and obtains the lens posture change sequence. The image pose matching module analyzes the video frame acquisition time and pose recording time based on the lens pose change sequence, filters adjacent pose records, adjusts the frame order and pose identifier pairing relationship, and obtains frame pose association identifier. The mucosal texture displacement module analyzes the changes in mucosal texture brightness arrangement in adjacent image frames based on the frame pose association identifier, compares the positions of the same texture point before and after, filters continuous texture points, and obtains the mucosal texture movement vector. The image rotation correction module compares the texture movement direction with the lens turning direction based on the mucosal texture movement vector, filters the rotation reference direction, and combines the cavity wall boundary with the image center position to obtain the viewing angle offset compensation amount. Based on the aforementioned viewpoint offset compensation amount, the lesion coordinate localization module analyzes the image chromaticity transition region, compares the brightness relationship between the inside and outside of the contour, filters continuous boundary pixels, adjusts the contour center relationship of adjacent image frames, and obtains the spatial contour coordinate data of the lesion.
2. The intelligent recognition and lesion localization system for digestive endoscopy images according to claim 1, characterized in that, The lens attitude change sequence includes lens pitch angle information, lens yaw angle information, and lens attitude time identifier. The frame attitude association identifier includes image frame number identifier, attitude time mapping identifier, and attitude matching index identifier. The mucosal texture movement vector includes texture point lateral displacement component, texture point longitudinal displacement component, and texture point movement direction identifier. The viewpoint offset compensation amount includes image rotation compensation angle, cavity wall center offset coordinates, and image center reference coordinates. The lesion spatial contour coordinate data includes lesion contour center coordinates, lesion boundary pixel coordinates, and lesion region spatial range.
3. The intelligent recognition and lesion localization system for digestive endoscopy images according to claim 1, characterized in that, The endoscope posture analysis module includes: The voltage timing recognition submodule is based on an electronic endoscope. It analyzes the sampling time corresponding to the voltage record of the angle detector of the endoscope body turning control part, compares the time sequence relationship of adjacent sampling segments, determines the continuous state of the time interval of the sampling segments, identifies the continuous time sequence segments, and obtains the continuous voltage timing segments. The bending direction determination submodule obtains voltage change direction records based on the voltage time sequence continuous segments, compares the correspondence between the displacement change direction of the turning steel wire and the voltage change direction, determines the records with consistent directions in the continuous segments, filters the set of segments with consistent directions, and obtains bending direction sequence data. The posture sequence construction submodule obtains the posture change information corresponding to the direction record based on the bending direction sequence data, compares the time position relationship of the upper and lower bending records, determines the time corresponding state of the left and right bending records, adjusts the time mapping relationship between the upper and lower bending records and the left and right bending records, and obtains the lens posture change sequence.
4. The intelligent recognition and lesion localization system for digestive endoscopy images according to claim 1, characterized in that, The image pose matching module includes: The timing correspondence submodule, based on the lens attitude change sequence, calls the lens video frame acquisition timing sequence and attitude recording timing, compares the sequential relationship of the timing in the same order, verifies the corresponding spacing of adjacent timings, and adjusts the positions of the alignable timings to obtain the frame attitude timing mapping identifier. The neighbor filtering submodule, based on the frame attitude time mapping identifier, calls the adjacent time interval records before and after, compares the relationship between the frame time and the attitude record spacing, determines the time adjacency order, merges the spacing continuous records, and obtains the attitude adjacency matching identifier. The frame order matching submodule calculates the attitude position corresponding to each frame based on the attitude adjacency matching identifier, combined with the image frame number and the arrangement order of attitude records, and corrects the corresponding order of frame order and attitude identifier to obtain the frame attitude association identifier.
5. The intelligent recognition and lesion localization system for digestive endoscopy images according to claim 1, characterized in that, The mucosal texture displacement module includes: The brightness sorting submodule, based on the frame attitude association identifier, calls the pixel position corresponding to the mucosal region of adjacent image frames, compares the brightness order of pixels at the same position, verifies the correspondence between the brightness transfer direction and the time sequence between frames, and consolidates the brightness sorting record of the region to obtain the pixel brightness order value. The texture coherence submodule, based on the pixel brightness sequence, calls the previous and subsequent position records of the same texture point, compares the overlap relationship between the forward displacement trajectory and the back-pointing position trajectory, determines the cross-frame movement coherence status of the texture point, filters texture points with consistent position connections, and obtains the texture trajectory coherence rate. The displacement synthesis submodule calculates the trajectory direction corresponding to the texture point combination direction based on the texture trajectory continuity rate, combines the lateral displacement record and the longitudinal displacement record, determines the belonging relationship of the turning position of the jump trajectory, adjusts the classification of abnormal turning trajectory, and obtains the mucosal texture movement vector.
6. The intelligent recognition and lesion localization system for digestive endoscopy images according to claim 1, characterized in that, The image rotation correction module includes: The same direction discrimination submodule obtains the texture movement direction and lens turning record based on the mucosal texture movement vector, compares the arrangement order of the direction marks in each frame, verifies the consistent pointing relationship in the same frame, filters out the corresponding reverse records, and obtains the rotation direction reference value. Based on the rotation direction reference, the boundary positioning submodule calls the image edge brightness transition trajectory, compares the connection order of adjacent transition points, determines the cavity wall boundary closure direction, and organizes the continuous boundary position records to obtain the cavity wall boundary positioning amount. The center correction submodule, based on the cavity wall boundary positioning amount, checks the relative position of the cavity wall center and the image center, compares the correspondence between the center offset direction and the rotation direction, adjusts the center correspondence order, and obtains the viewpoint offset compensation amount.
7. The intelligent recognition and lesion localization system for digestive endoscopy images according to claim 1, characterized in that, The lesion coordinate localization module includes: Based on the aforementioned viewpoint offset compensation amount, the transition distribution submodule calls the colorimetric record of the digestive tract mucosal wall structure image, compares the colorimetric change order of adjacent pixels with the brightness transition relationship between the inner and outer sides of the contour, and adjusts the position of the continuous transition area to obtain the colorimetric transition distribution rate. Based on the chromaticity transition distribution rate, the boundary path submodule calls the boundary pixel adjacency record, compares the connection order of adjacent boundary pixels with the consistency of path turning, filters out continuous closed path segments, and straightens the connection direction to obtain the boundary closure path rate. Based on the boundary closure path rate, the contour connection submodule checks the corresponding position after the boundary coordinates are rotated, compares the displacement relationship between the contour centers of adjacent image frames, and adjusts the connection relationship between the contour centers of adjacent image frames to obtain the spatial contour coordinate data of the lesion.
8. The intelligent recognition and lesion localization system for digestive endoscopy images according to claim 1, characterized in that, The sampling time refers to the time marker corresponding to when the angle detector collects the voltage signal. The vertical bending refers to the voltage sampling record when the lens is bent in the vertical direction, and the horizontal bending refers to the voltage sampling record when the lens is bent in the horizontal direction.
9. The intelligent recognition and lesion localization system for digestive endoscopy images according to claim 1, characterized in that, The video frame acquisition time refers to the time sequence formed by arranging the time markers corresponding to each image frame in the electronic endoscope video stream in chronological order when they are acquired, and the attitude recording time refers to the time marker corresponding to the lens attitude angle recording during the sampling process.
10. The intelligent recognition and lesion localization system for digestive endoscopy images according to claim 1, characterized in that, The pairing relationship refers to the correspondence established between image frame acquisition records and posture records according to the time sequence, and the preceding and following positions refer to the spatial correspondence between the pixel positions corresponding to the same texture point in two consecutive image frames.