Image recognition system and method based on deep learning
By combining texture features and deep learning morphological prediction models in low-frame medical images, the problem of cross-frame association of surgical instruments is solved, and stable tracking and accurate identification of surgical instruments is achieved, which improves the continuity and robustness of recognition.
Patent Information
- Application Number
- CN202510949759.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-07-10
AI Technical Summary
In low-frame rate medical images, existing deep learning image recognition methods are difficult to stably associate surgical instruments between adjacent frames, resulting in lost or misidentified target identity, affecting the accuracy of instrument trajectory reconstruction and intraoperative navigation.
By obtaining texture data of time-series adjacent image frames in medical images, combining texture feature positioning and deep learning morphological prediction models, predicting the target evolution state, and achieving cross-frame target correlation through feature matching, dynamically adjusting the matching threshold to adapt to image quality changes.
In low-frame-rate medical images, stable tracking and accurate identification of surgical instruments are achieved, which improves the continuity and robustness of target recognition, reduces recognition interruptions and misjudgments, and improves the reliability of image-assisted analysis and navigation.
Smart Images

Figure CN120451519A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an image recognition system and method based on deep learning. Background Art
[0002] Image recognition technology started in the 1990s and has evolved from traditional computer vision algorithms based on edges and textures to statistical learning methods and then to deep learning models. Especially after the emergence of deep neural networks, architectures such as convolutional neural networks, region proposal networks, YOLO and Transformer have promoted the rapid development of target detection and image understanding, and are widely used in scenarios such as autonomous driving, video surveillance, and intelligent manufacturing. Today, with the continuous improvement of algorithm accuracy, research trends are gradually shifting to low-latency recognition, edge device deployment, and enhanced recognition robustness in complex dynamic environments.
[0003] Existing deep learning image recognition methods have been able to achieve high recognition accuracy by learning spatial features in static frames. However, for medical images, such as laparoscopic minimally invasive surgery images, which are usually collected and recorded by endoscopic cameras for purposes such as intraoperative auxiliary recognition, postoperative behavior analysis, and clinical teaching, these images are often set to a low frame rate during the acquisition or processing process due to factors such as long surgery time, high image storage pressure, high manual annotation costs, and remote transmission bandwidth limitations. Some are even provided at 1fps in public research datasets. Compared with high-frame-rate videos, the time interval between adjacent frames in low-frame-rate videos is larger, resulting in significant appearance changes such as position offset, angular rotation, or deformation of surgical instruments between previous and next frames. This makes it difficult for traditional image recognition methods based on spatial feature consistency to stably extract target features or accurately complete the association of previous and next frames, and is prone to target identity loss, jumps, or misidentification. This recognition discontinuity will directly affect the stability and accuracy of key applications such as instrument trajectory reconstruction, surgical stage identification, and intraoperative navigation.
[0004] Therefore, an image recognition system and method based on deep learning are proposed. Summary of the Invention
[0005] In view of the above-mentioned state of the art, the present application is proposed. The embodiments of the present application provide an image recognition system and method based on deep learning, which can improve the accuracy of continuous recognition of surgical instruments in low-frame-rate medical images.
[0006] According to one aspect of the present application, a deep learning-based image recognition method is provided, comprising: acquiring a first image frame and a second image frame in a medical image that are adjacent in temporal position and have a time interval greater than a preset interval threshold; acquiring first texture data of a jaw area corresponding to a surgical instrument in the medical image; locating a first target in the first image frame based on the first texture data; locating at least one candidate target in the second image frame, the candidate target being a possible continuation target of the first target in the second image frame; extracting a first feature vector of the first target in the first image frame and at least one previous image frame to form a first feature vector set; extracting a second feature vector of each candidate target; generating a third feature vector representing the evolution state of the first target between the first image frame and the second image frame based on the first feature vector set through a deep learning-based morphological prediction model; matching the third feature vector with each of the second feature vectors to determine whether there is a second feature vector with the highest matching degree and exceeding a preset matching threshold, and if so, determining the candidate target corresponding to the second feature vector as the continuation target of the first target in the second image frame, otherwise determining that the first target is lost in the second image frame.
[0007] According to another aspect of the present application, a deep learning-based image recognition system is provided, including: an image frame acquisition module for acquiring a first image frame and a second image frame in a medical image that are adjacent in temporal position and have a time interval greater than a preset interval threshold; a texture data acquisition module for acquiring first texture data of a jaw area corresponding to a surgical instrument in the medical image; a first target positioning module for positioning a first target in the first image frame based on the first texture data; a candidate target positioning module for positioning at least one candidate target in the second image frame, the candidate target being a possible continuation target of the first target in the second image frame; and a historical feature extraction module for extracting the first target in the first image frame and at least one previous image frame. The first feature vector in the frame constitutes a first feature vector set; a candidate feature extraction module is used to extract the second feature vector of each candidate target; a morphology prediction module is used to generate a third feature vector representing the evolution state of the first target between the first image frame and the second image frame according to the first feature vector set through a morphology prediction model based on deep learning; a matching decision module is used to match the third feature vector with each of the second feature vectors to determine whether there is a second feature vector with the highest matching degree and exceeding a preset matching threshold. If so, the candidate target corresponding to the second feature vector is determined as the continuation target of the first target in the second image frame; otherwise, the first target is determined to be lost in the second image frame.
[0008] According to another aspect of the present application, an electronic device is provided, comprising a memory and a processor, wherein the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which implement the steps of the above-described method when executed by the processor.
[0009] According to another aspect of the present application, a computer storage medium is provided, on which computer executable instructions are stored. When the computer executable instructions are executed by a processor, the steps of the above method are implemented.
[0010] Compared with the existing technology, the deep learning-based image recognition system and method according to the embodiments of the present application can achieve stable tracking and accurate identification of surgical instruments in low-frame-rate medical images, improve the continuity and robustness of target recognition, and effectively reduce recognition interruptions or misjudgments caused by image blur, insufficient frame rate or target occlusion, thereby improving the reliability of image-assisted analysis and navigation. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0012] Figure 1 This is a flowchart of the image recognition method based on deep learning of the present invention.
[0013] Figure 2 This is a flowchart of the secondary search continuation target of the deep learning-based image recognition method of the present invention.
[0014] Figure 3 This is a flowchart of the dynamic adjustment of the preset matching threshold of the deep learning-based image recognition method of the present invention.
[0015] Figure 4 This is a block diagram of the image recognition system based on deep learning of the present invention.
[0016] Figure 5 The present invention is a block diagram of an electronic device. DETAILED DESCRIPTION
[0017] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0018] Application Overview
[0019] In the existing technology, image recognition technology has evolved from traditional computer vision algorithms to deep learning models, but there are still significant challenges in low-frame-rate medical imaging scenarios. For low-frame-rate medical images, the large interval between adjacent frames may cause significant positional offset, angular rotation or deformation of surgical instruments. Traditional recognition methods based on spatial feature consistency are difficult to effectively associate targets in previous and next frames, resulting in target identity loss or misidentification, affecting the accuracy of instrument trajectory reconstruction and intraoperative navigation.
[0020] In the existing technology, the core of the difficulty in associating targets in low-frame-rate images lies in the lack of inter-frame states and insufficient prediction of morphological changes. Through analysis, it is found that the texture of key parts of instruments is relatively stable and can be used as the basis for initial positioning. At the same time, the target's motion trajectory in time is continuous, and the evolution trend can be predicted through historical feature modeling. Based on this, the basic idea of this application is to propose a technical route that combines texture feature positioning, multi-frame feature accumulation and deep learning prediction model, and solve the problem of inter-frame association by predicting the target evolution state and dynamically matching it with candidate features.
[0021] Exemplary Methods
[0022] Figures 1 to 3 The diagram illustrates an image recognition method based on deep learning according to an embodiment of the present application, including: obtaining a first image frame and a second image frame in a medical image whose temporal positions are adjacent and whose time interval is greater than a preset interval threshold; obtaining first texture data of the jaw area corresponding to the surgical instrument in the medical image; locating a first target in the first image frame according to the first texture data; locating at least one candidate target in the second image frame, the candidate target being a possible continuation target of the first target in the second image frame; extracting a first feature vector of the first target in the first image frame and at least one previous image frame to form a first feature vector set; extracting a second feature vector of each candidate target; generating a third feature vector representing the evolution state of the first target between the first image frame and the second image frame according to the first feature vector set through a morphological prediction model based on deep learning; matching the third feature vector with each second feature vector to determine whether there is a second feature vector with the highest matching degree and exceeding the preset matching threshold. If so, the candidate target corresponding to the second feature vector is determined as the continuation target of the first target in the second image frame, otherwise it is determined that the first target is lost in the second image frame.
[0023] The first texture data refers to the microscopic texture feature set of the surgical instrument jaw surface. In the application, obtaining the first texture data can be achieved by matching it with a pre-built database of surgical instrument texture feature sets. The morphology prediction model refers to a neural network capable of learning the temporal evolution of the target morphology. Specifically, it can be implemented using a long short-term memory network with an attention mechanism. It can predict the morphological change trend of the target during the frame interval by analyzing the historical feature sequence.
[0024] Specifically, this method uses a preset interval threshold to filter out key frame pairs with a sufficiently large time span to avoid processing redundant frame data. The unique texture features of the instrument jaw area are used for initial target positioning to ensure basic detection accuracy. After establishing a set of candidate targets in subsequent frames, a feature set containing temporal evolution information is constructed by extracting the deep features of the target in the current frame and multiple historical frames. The morphological prediction model performs temporal modeling on this feature set to generate predicted features that reflect the possible evolutionary state of the target. By matching the predicted features with the real-time features of the candidate targets for similarity, accurate association of targets across frames is ultimately achieved.
[0025] Through the above technical solution, this application effectively solves the problem of difficulty in cross-frame association of surgical instruments in low-frame-rate medical images. By combining texture feature positioning with multi-frame feature prediction, while ensuring the initial detection accuracy, it improves the accuracy of cross-frame target matching, ensures the continuity of the instrument recognition process, and provides a reliable data foundation for subsequent surgical navigation and behavior analysis.
[0026] The present application further proposes that before determining that the first target is lost in the second image frame, it also includes: determining a predicted candidate area for the first target in the second image frame based on the third eigenvector, the predicted candidate area being an image area of a preset size centered on the position corresponding to the third eigenvector; extracting multiple new candidate targets within the predicted candidate area, and extracting the fourth eigenvector of each new candidate target respectively; matching the third eigenvector with each fourth eigenvector, and if there is a fourth eigenvector with the highest matching degree and exceeding the preset secondary threshold, then determining the new candidate target corresponding to the fourth eigenvector as the continuation target of the first target in the second image frame.
[0027] Among them, the predicted candidate area refers to the candidate search range generated based on the spatial position corresponding to the third eigenvector output by the morphological prediction model. Specifically, it can be implemented by a rectangular area with the predicted coordinates as the center and the side length as the preset pixel value, which is used to cover the possible displacement range of the device in the low frame rate image. The new candidate target refers to the potential device area extracted by the sliding window or region proposal network within the predicted candidate area. Specifically, it can be implemented by using a multi-scale sliding window combined with an edge detection algorithm to generate a candidate frame to supplement the omissions of the original candidate target set. The preset secondary threshold refers to the feature similarity judgment standard that is lower than the preset matching threshold. Specifically, it can be implemented by a fixed value of 0.8 times the main matching threshold or a value dynamically adjusted based on historical matching data, which is used to relax the matching conditions within the extended search range.
[0028] Specifically, when the initial candidate target fails to match, an image area covering the possible displacement range of the instrument is generated based on the spatial coordinates corresponding to the third eigenvector output by the morphological prediction model. By performing dense sampling or region proposal operations within this area, multiple new candidate targets are extracted and their depth features are extracted. The third eigenvector is calculated for similarity with the depth features of the new candidate target. When the highest match exceeds the secondary threshold, the candidate target is associated with a continued instance of the same instrument. For example, when the instrument rotates rapidly in a low-frame-rate image, resulting in a mismatch in the original candidate target features, the candidate area is predicted to cover the new position of the instrument after rotation, and the deformed instrument instance is captured through the secondary matching mechanism.
[0029] Through the above technical solution, the present application improves the continuous recognition capability of instrument targets in low-frame-rate medical images by expanding the candidate area and matching it again when the instrument moves rapidly or undergoes sudden morphological changes, and further reduces target association failures caused by excessive changes between frames.
[0030] The present application further proposes that before determining whether the matching degree exceeds the preset matching threshold, it also includes: calculating the average clarity of the first image frame and the second image frame; determining the average clarity: if it is lower than the first clarity threshold, setting the preset matching threshold to the first matching threshold; if it is not lower than the first clarity threshold and lower than the second clarity threshold, setting the preset matching threshold to the second matching threshold; if it is not lower than the second clarity threshold, setting the preset matching threshold to the third matching threshold; wherein, the values of the first matching threshold, the second matching threshold and the third matching threshold are incremented.
[0031] Among them, average clarity refers to an indicator that quantifies the degree of image detail retention by calculating the statistical average of the pixel gradient amplitudes in the image frame. Specifically, it can be achieved by using the Sobel operator to extract the image edge and then calculating the gradient amplitude mean. This indicator can reflect the degree of image blur. The first clarity threshold and the second clarity threshold refer to the critical values of image quality grading that have been pre-calibrated through experiments. Specifically, they can be determined by statistically calculating the inflection points of feature matching accuracy at different clarity levels on a standard test set, and are used to divide the image quality into three levels: low, medium, and high. The preset matching threshold refers to the critical value for judging the similarity of feature vectors. Specifically, it can be achieved by using the confidence value of cosine similarity or Euclidean distance transformation. Its numerical setting directly affects the strictness of the matching results.
[0032] Specifically, in the process of low-frame-rate medical image processing, the gradient amplitude of each pixel in the two frames of images is calculated separately by the Sobel operator, and the average value of the gradient amplitude of all pixels is taken as the average clarity. When the value is lower than the pre-calibrated first clarity threshold, it indicates that the image is obviously blurred. At this time, the reliability of feature extraction decreases, and the matching threshold is lowered to the first matching threshold to avoid missed recognition due to feature deviation. When the average clarity is in the middle range, a medium-strict second matching threshold is used to balance the recognition accuracy and recall rate. When the image quality reaches the high-definition standard, a higher third matching threshold is used to suppress false matches. This hierarchical control mechanism establishes a positive correlation between image quality and matching threshold, which effectively alleviates the problem of missed detection in low-quality images while ensuring recognition accuracy in high-definition scenes.
[0033] Through the above technical solution, the present application can maintain continuous tracking of instrument targets by adaptively lowering matching requirements when intermittent quality degradation occurs in low-frame-rate medical images, and automatically improve matching standards to ensure recognition accuracy when image quality is restored, effectively solving the problem of misidentification of instruments caused by fluctuations in image clarity.
[0034] The present application further proposes that after determining the average clarity, it also includes: extracting the area other than the first target in the first image frame as the first background area; extracting the area other than at least one candidate target area in the second image frame as the second background area; calculating the average texture complexity of the first background area and the second background area; determining whether the average texture complexity is higher than a preset complexity threshold, and if so, increasing the preset matching threshold according to the difference between the average texture complexity and the preset complexity threshold through a preset first adjustment function.
[0035] Among them, the first background area and the second background area refer to the non-target areas remaining after the identified instrument target area and the suspected instrument target area in the first image frame are excluded by image segmentation technology. Specifically, this can be achieved by a region division method based on edge detection, which is used to isolate the interference of the instrument target on the background texture analysis. The average texture complexity refers to the texture feature statistics of the background area obtained by calculating the gray-level co-occurrence matrix. Specifically, it can be achieved by a weighted combination index of energy, contrast, and entropy value, which is used to quantify the distribution density of textures similar to the instrument target in the background. The preset complexity threshold refers to the critical value of the background interference intensity obtained by training historical data. Specifically, it can be achieved by the probability threshold output by the machine learning classifier, which is used to determine whether it is necessary to start the matching threshold adjustment. The first adjustment function refers to the mathematical relationship for dynamically adjusting the matching threshold according to the difference between the background complexity and the threshold. Specifically, it can be achieved by a linear proportional function or an exponential function, which is used to establish a positive correlation between the background interference intensity and the matching strictness.
[0036] Specifically, the non-target areas in the two frames before and after are first extracted as background analysis objects through a region segmentation method. Subsequently, a texture feature extraction algorithm is used to quantitatively evaluate the two background areas, and a comprehensive index reflecting the complexity of the texture is calculated. When this index exceeds a preset threshold, it indicates that there are a large number of interfering features in the background that are similar to the texture of the instrument jaws. At this time, a predefined adjustment function is used to increase the matching threshold requirement based on the degree of deviation between the actual complexity and the threshold, so that the candidate target must have a higher similarity with the predicted feature to be judged as a continued target. This dynamic adjustment can avoid mismatches caused by similar textures in complex backgrounds, while retaining the possibility of matching the real target within a reasonable deformation range.
[0037] Through the above technical solution, the present application can dynamically optimize the matching threshold setting according to the background texture characteristics of the actual surgical scene, effectively distinguish between real instrument movement and background interference characteristics in low-frame-rate medical images with complex tissue textures, reduce the probability of false matching due to similar textures, and avoid overly strict threshold settings that cause missed matching of real targets, further ensuring the continuous tracking stability of instrument targets between previous and subsequent frames.
[0038] This application further proposes:
[0039] Locating the first target in the first image frame according to the first texture data includes: calculating a first similarity between each image region in the first image frame and the first texture data; determining an image region having the highest first similarity and exceeding a preset similarity threshold as the first target;
[0040] Positioning at least one candidate target in the second image frame includes: calculating a second similarity between each image region in the second image frame and the first texture data; and selecting an image region whose second similarity exceeds a preset similarity threshold as a candidate target.
[0041] The first and second similarities refer to the degree of match between an image region and a standard texture feature. These can be achieved using a structural similarity index or cosine similarity calculation using convolutional neural network features to quantify the relevance of the region to the target texture. The preset similarity threshold is the minimum similarity requirement for selecting valid matching regions.
[0042] The present application further proposes extracting the first feature vector of the first target in the first image frame and at least one previous image frame to obtain a first feature vector set, including: obtaining at least one image frame before the first image frame; extracting the position information of the first target in the first image frame; predicting the possible position area of the first target in at least one image frame based on the position information; judging whether there is an associated image area in at least one image frame that is located in the possible position area and has a similarity with the first texture data exceeding a preset judgment threshold; extracting the first feature vectors of each associated image area and the first target to form a first feature vector set.
[0043] Among them, the at least one previous image frame refers to several historical frame data arranged in time sequence before the current frame. Specifically, it can be implemented by selecting the latest N frames of images using a sliding window mechanism to expand the feature source of the time dimension. Position information refers to the spatial coordinates and area range of the target in the current frame. The possible position area refers to the existence range of the historical frame target based on the current frame position. Specifically, the motion trajectory extrapolation algorithm can be used to generate a rectangular area to limit the search range of the historical frame. The associated image area refers to the image block in the historical frame that meets the spatial constraints and texture matching. Specifically, it can be screened by combining regional segmentation and similarity calculation to ensure the consistency of historical features with the current target.
[0044] Specifically, during low-frame-rate medical image processing, a sliding window is used to obtain several historical frames before the current frame. The target's detection frame coordinates in the current frame are used to construct a motion trajectory, and the rectangular area where the target may exist in the historical frame is inferred. For each historical frame, image segmentation is performed within the predicted area, and the similarity between each sub-region and the preset jaw texture is calculated, and candidate regions exceeding the threshold are screened out. The feature vectors of these candidate regions are aggregated with the target features of the current frame to form a data set containing time series features. This process uses a dual-time and space constraint mechanism to not only narrow the search range by using motion continuity, but also eliminate interference areas by texture similarity, effectively integrating multi-frame feature information.
[0045] Exemplary Systems
[0046] Figure 4 The diagram shows an image recognition system based on deep learning according to an embodiment of the present application, including: an image frame acquisition module for acquiring a first image frame and a second image frame in a medical image whose temporal positions are adjacent and whose time interval is greater than a preset interval threshold; a texture data acquisition module for acquiring first texture data of the jaw area corresponding to the surgical instrument in the medical image; a first target positioning module for locating a first target in the first image frame according to the first texture data; a candidate target positioning module for locating at least one candidate target in the second image frame, the candidate target being a possible continuation target of the first target in the second image frame; a historical feature extraction module for extracting the first target in the first image frame and at least one previous image frame. The first feature vector in the frame constitutes a first feature vector set; a candidate feature extraction module is used to extract the second feature vector of each candidate target; a morphology prediction module is used to generate a third feature vector representing the evolution state of the first target between the first image frame and the second image frame according to the first feature vector set through a morphology prediction model based on deep learning; a matching decision module is used to match the third feature vector with each second feature vector to determine whether there is a second feature vector with the highest matching degree and exceeding a preset matching threshold. If so, the candidate target corresponding to the second feature vector is determined as the continuation target of the first target in the second image frame; otherwise, the first target is determined to be lost in the second image frame.
[0047] In one example, the matching decision module determines that before the first target is lost in the second image frame, it also includes: determining a predicted candidate area for the first target in the second image frame based on the third eigenvector, the predicted candidate area being an image area of a preset size centered on the position corresponding to the third eigenvector; extracting multiple new candidate targets within the predicted candidate area, and extracting the fourth eigenvector of each new candidate target respectively; matching the third eigenvector with each fourth eigenvector, and if there is a fourth eigenvector with the highest matching degree and exceeding the preset secondary threshold, then determining the new candidate target corresponding to the fourth eigenvector as the continuation target of the first target in the second image frame.
[0048] In one example, before the matching decision module determines whether the matching degree exceeds the preset matching threshold, it also includes: calculating the average clarity of the first image frame and the second image frame; judging the average clarity: if it is lower than the first clarity threshold, setting the preset matching threshold to the first matching threshold; if it is not lower than the first clarity threshold and lower than the second clarity threshold, setting the preset matching threshold to the second matching threshold; if it is not lower than the second clarity threshold, setting the preset matching threshold to the third matching threshold; wherein the values of the first matching threshold, the second matching threshold and the third matching threshold are incremented.
[0049] In one example, after determining the average clarity, the matching decision module further includes: extracting an area other than the first target in the first image frame as a first background area; extracting an area other than at least one candidate target area in the second image frame as a second background area; calculating the average texture complexity of the first background area and the second background area; determining whether the average texture complexity is higher than a preset complexity threshold, and if so, increasing the preset matching threshold according to the difference between the average texture complexity and the preset complexity threshold through a preset first adjustment function.
[0050] In one example, the first target positioning module positions the first target in the first image frame based on the first texture data, including: calculating the first similarity between each image area in the first image frame and the first texture data; and determining the image area with the highest first similarity and exceeding a preset similarity threshold as the first target.
[0051] In one example, the candidate target positioning module positions at least one candidate target in the second image frame, including: calculating a second similarity between each image region in the second image frame and the first texture data; and selecting an image region where the second similarity exceeds a preset similarity threshold as a candidate target.
[0052] In one example, a historical feature extraction module extracts a first feature vector of a first target in a first image frame and at least one previous image frame, and the first feature vector set is formed by: obtaining at least one image frame before the first image frame; extracting position information of the first target in the first image frame; predicting a possible position area of the first target in at least one image frame based on the position information; determining whether there is an associated image area in at least one image frame that is located within the possible position area and has a similarity with the first texture data exceeding a preset judgment threshold; extracting the first feature vectors of each associated image area and the first target to form a first feature vector set.
[0053] Exemplary electronic devices
[0054] Figure 5 The figure shows an electronic device according to an embodiment of the present application. The electronic device can be the mobile device itself, or a stand-alone device independent of the mobile device, which can communicate with the mobile device to receive collected input signals from the mobile device and send the selected target driving behavior to the mobile device.
[0055] Figure 5 The figure shows a block diagram of an electronic device according to an embodiment of the present application.
[0056] like Figure 5 As shown, the electronic device includes one or more processors and memory.
[0057] The processor may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device to perform desired functions.
[0058] The memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor may execute the program instructions to implement the driving behavior decision-making method of each embodiment of the present application described above and / or other desired functions.
[0059] In one example, the electronic device may further include an input device and an output device, and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0060] Of course, to simplify, Figure 5 Only some of the components in the electronic device related to the present application are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, the electronic device may further include any other appropriate components according to specific application scenarios.
[0061] Exemplary computer-readable media
[0062] An embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, causes the processor to execute the steps of the driving behavior decision-making method according to various embodiments of the present application described in the above “Exemplary Method” section of this specification.
[0063] Computer-readable storage media can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or components, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0064] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this application are merely illustrative and not restrictive, and it should not be assumed that these advantages, strengths, and effects are required of each embodiment of this application. In addition, the specific details disclosed above are merely illustrative and facilitating understanding, and are not restrictive. The above details do not limit this application to necessarily being implemented using the above specific details.
[0065] The block diagrams of the devices, devices, equipment, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0066] It should also be noted that in the apparatus, device, and method of the present application, each component or each step can be decomposed and / or recombined, and such decomposition and / or recombination should be regarded as equivalent solutions of the present application.
[0067] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present application. Therefore, the present application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0068] The above description has been provided for the purpose of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present application to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A deep learning-based image recognition method is applied to device recognition in low-frame-rate medical images, characterized by: include: Acquire a first image frame and a second image frame in a medical image that are adjacent in time sequence and whose time interval is greater than a preset interval threshold; Acquire first texture data of a jaw region corresponding to the surgical instrument in the medical image; locating a first target in the first image frame according to the first texture data; Locating at least one candidate target in the second image frame, where the candidate target is a possible continuation target of the first target in the second image frame; extracting a first feature vector of the first target in the first image frame and at least one previous image frame to form a first feature vector set; Extracting a second feature vector of each candidate target; generating, by a deep learning-based morphology prediction model, a third feature vector representing an evolution state of the first object between the first image frame and the second image frame according to the first feature vector set; Match the third eigenvector with each of the second eigenvectors to determine whether there is a second eigenvector with the highest matching degree and exceeding a preset matching threshold; if so, determine the candidate target corresponding to the second eigenvector as the continuation target of the first target in the second image frame; otherwise, determine that the first target is lost in the second image frame.
2. The image recognition method based on deep learning according to claim 1, characterized in that The determining that the first target is lost in the second image frame further includes: determining a prediction candidate region of the first target in the second image frame according to the third eigenvector, the prediction candidate region being an image region of a preset size centered at a position corresponding to the third eigenvector; Extracting a plurality of new candidate targets within the prediction candidate area, and respectively extracting a fourth eigenvector of each of the new candidate targets; The third eigenvector is matched with each of the fourth eigenvectors. If there is a fourth eigenvector with the highest matching degree and exceeding a preset secondary threshold, the new candidate target corresponding to the fourth eigenvector is determined as the continuation target of the first target in the second image frame.
3. The image recognition method based on deep learning according to claim 1, characterized in that Before determining whether the matching degree exceeds a preset matching threshold, the method further includes: Calculating an average clarity of the first image frame and the second image frame; Determine the average clarity: If the clarity is lower than the first clarity threshold, setting the preset matching threshold as the first matching threshold; If the value is not lower than the first clarity threshold and lower than the second clarity threshold, setting the preset matching threshold as the second matching threshold; If it is not lower than the second clarity threshold, setting the preset matching threshold to the third matching threshold; The first matching threshold, the second matching threshold and the third matching threshold are numerically increased.
4. The image recognition method based on deep learning according to claim 3, characterized in that After determining the average clarity, the method further includes: extracting an area excluding the first target in the first image frame as a first background area; extracting an area other than the at least one candidate target area in the second image frame as a second background area; Calculating average texture complexity of the first background area and the second background area; It is determined whether the average texture complexity is higher than a preset complexity threshold. If so, the preset matching threshold is increased according to the difference between the average texture complexity and the preset complexity threshold by using a preset first adjustment function.
5. The image recognition method based on deep learning according to claim 1, characterized in that The locating the first target in the first image frame according to the first texture data includes: calculating a first similarity between each image region in the first image frame and the first texture data; An image region with the highest first similarity and exceeding a preset similarity threshold is determined as the first target.
6. The image recognition method based on deep learning according to claim 5, characterized in that: The locating at least one candidate target in the second image frame comprises: calculating a second similarity between each image region in the second image frame and the first texture data; The image region where the second similarity exceeds the preset similarity threshold is taken as a candidate target.
7. The image recognition method based on deep learning according to claim 1, characterized in that: Extracting a first feature vector of the first target in the first image frame and at least one previous image frame to form a first feature vector set includes: Acquire at least one image frame preceding the first image frame; Extracting position information of the first target in the first image frame; predicting a possible location area of the first target in the at least one image frame according to the location information; Determining whether there is an associated image region in the at least one image frame that is located within the possible position area and has a similarity with the first texture data exceeding a preset determination threshold; The first feature vectors of each of the associated image regions and the first target are extracted to form the first feature vector set.
8. An image recognition system based on deep learning, characterized in that: include: An image frame acquisition module, configured to acquire a first image frame and a second image frame in a medical image that are adjacent in time sequence and have a time interval greater than a preset interval threshold; a texture data acquisition module, configured to acquire first texture data of a jaw region corresponding to the surgical instrument in the medical image; a first target positioning module, configured to locate a first target in the first image frame according to the first texture data; a candidate target positioning module, configured to locate at least one candidate target in the second image frame, where the candidate target is a possible continuation target of the first target in the second image frame; a historical feature extraction module, configured to extract a first feature vector of the first target in the first image frame and at least one previous image frame to form a first feature vector set; A candidate feature extraction module, configured to extract a second feature vector of each candidate target; a morphology prediction module, configured to generate, based on the first set of feature vectors, a third feature vector representing an evolution state of the first target between the first image frame and the second image frame, using a deep learning-based morphology prediction model; A matching decision module is used to match the third eigenvector with each of the second eigenvectors to determine whether there is a second eigenvector with the highest matching degree and exceeding a preset matching threshold. If so, the candidate target corresponding to the second eigenvector is determined as the continuation target of the first target in the second image frame; otherwise, the first target is determined to be lost in the second image frame.
9. An electronic device comprising a memory and a processor, characterized in that: The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer storage medium having computer-executable instructions stored thereon, characterized in that: When the computer-executable instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Microsurgery auxiliary system based on visual identification
CN120072216A
Systems and methods for robotic medical system integration with external imaging
US12161434B2
Systems for detecting and tracking of objects and co-registration
WO2014124447A1