Marker Recognition Method, Device, Equipment and Storage Medium
By processing and feature extraction of video frames collected by the endoscopy, and combining reference object information, the real size information of the lesion is calculated, the problem of inaccurate lesion size recognition in endoscopy is solved and the accuracy of diagnosis is improved.
Patent Information
- Application Number
- CN202410043794.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-01-11
AI Technical Summary
In digestive tract endoscopy, the identification of lesions is not accurate enough, resulting in inaccurate diagnosis of the disease.
By acquiring the video frames collected by the endoscopy, the first image is obtained, and the marker size information and reference size information are extracted for feature extraction, the video frame is analyzed to determine the distance value between the marker and the endoscopy, and the real size information of the target marker is calculated based on this information.
It improves the accuracy of identification of markers such as lesions, reduces the impact of human subjective judgment, and enhances the reliability of disease diagnosis.
Smart Images

Figure CN117876326B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of medical assistance technologies, and particularly to a method and device for identifying markers, an electronic device, and a storage medium. Background Art
[0002] An endoscope is a medical instrument that enters the human body through a tube to observe the internal conditions of the human body. Endoscopic examination can achieve the purpose of observing the internal organs of the human body with the least harm and is a very important means of observation and treatment in modern medicine.
[0003] During the process of digestive tract endoscopic examination, it is usually necessary to measure the size of the lesion, and then judge the disease risk level according to the size of the lesion. At the same time, the treatment method for the disease can also be determined. In clinical practice, it is usually judged by an endoscopist with the naked eye. Due to the large subjectivity and the lack of specific obvious references, the identification and determination of the lesion size are not accurate enough, which will lead to inaccurate disease diagnosis. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a method and device for identifying markers, an electronic device, and a storage medium to solve the technical problem of inaccurate identification of the size of markers such as lesions in the related art.
[0005] In a first aspect, the embodiments of the present application provide a method for identifying markers, including:
[0006] Obtain a video frame collected by an endoscope, and process the video frame to obtain a first image;
[0007] Extract features from the first image to obtain the marker size information and reference object size information corresponding to the first image;
[0008] Analyze the video frame to obtain the distance value between the endoscope and the target marker corresponding to the marker size information;
[0009] According to the distance value, the marker size information, and the reference object size information, obtain the true size information of the target marker.
[0010] In a second aspect, the embodiments of the present application provide a device for identifying markers, including:
[0011] A video acquisition module, configured to obtain a video frame collected by an endoscope, and process the video frame to obtain a first image;
[0012] A feature extraction module, configured to extract features from the first image to obtain the marker size information and reference object size information corresponding to the first image;
[0013] A video analysis module, configured to analyze the video frame to obtain a distance value of a target marker corresponding to the size information of the endoscope and the marker;
[0014] A result processing module, configured to obtain true size information of the target marker according to the distance value, the marker size information, and the reference object size information.
[0015] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in any one of the above-mentioned marker recognition methods are implemented.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in any one of the above-mentioned marker recognition methods are implemented.
[0017] An embodiment of the present application provides a marker recognition method, device, electronic device, and storage medium. In the process of recognizing and processing markers such as lesions, video frames collected by an endoscope are obtained, and the collected video frames are screened and processed to obtain a first image for marker recognition. Then, feature information of the first image is extracted to obtain marker size information and reference object size information corresponding to the first image. At the same time, by analyzing the video frames, the distance relationship between the currently recognized marker and the endoscope is determined, such as a distance value. Furthermore, according to the obtained marker size information, reference object size information, and distance relationship, the true size information of the target marker included in the first image is determined. It realizes that in the recognition and processing process, a reference object is added, the markers and reference objects included in the collected video frames are recognized, and then the true size information of the markers is obtained by using the set position information of the reference object, etc., improving the accuracy of recognizing markers such as lesions. Description of the Drawings
[0018] Figure 1 is a flowchart of a marker recognition method provided by an embodiment of the present application;
[0019] Figure 2 is a flowchart of obtaining a first image provided by an embodiment of the present application;
[0020] Figure 3 is a schematic diagram of cropping a video frame provided by an embodiment of the present application;
[0021] Figure 4 is a schematic diagram of scaling a cropped video frame provided by an embodiment of the present application;
[0022] Figure 5 It is a schematic diagram of pixel adjustment for scaling the cropped video frame provided by an embodiment of the present application;
[0023] Figure 6 It is a flowchart of the steps for obtaining the size information of the reference object provided by an embodiment of the present application;
[0024] Figure 7 It is a flowchart of the steps for obtaining the size information of the marker provided by an embodiment of the present application;
[0025] Figure 8 It is a schematic structural diagram of a marker recognition device provided by an embodiment of the present application;
[0026] Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;
[0027] Figure 10 It is another schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0029] It should be understood that the various steps recorded in the method embodiments of the present disclosure may be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0030] The term "including" and its variants used herein are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0031] In the related art, during gastrointestinal endoscopy, it is usually necessary to measure the size of a lesion, and then judge the disease risk level according to the size of the lesion. At the same time, the treatment method for the disease can also be determined. However, in the actual processing, it is usually the endoscopist who makes a visual judgment based on experience, resulting in inaccurate identification of the lesion size, and thus may lead to inaccurate disease diagnosis.
[0032] To solve the technical problems existing in the related art, an embodiment of the present application provides a marker recognition method. Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a marker recognition method provided by an embodiment of the present application. The method includes steps 101 to 104.
[0033] Step 101, obtain a video frame collected by an endoscope, and process the video frame to obtain a first image.
[0034] In one embodiment, obtain the video frame collected by the endoscope, and perform corresponding processing on the video frame to obtain a first image for analyzing markers such as lesions. Then, through the analysis and processing of the first image, determine the true size information of the markers included.
[0035] Exemplarily, when the endoscope is in the human body's pipeline, control the corresponding image acquisition device to obtain a video, and obtain the corresponding video frame. When determining the relevant information of the patient's lesions and other markers, such as the lesion location and lesion size, it is necessary to perform corresponding processing on the video frame, such as determining which video frames contain lesions and the size information of the included lesions. Through the acquisition of the video frame, obtain an image that can be used to obtain the relevant feature information of the lesions and other markers, and then perform analysis and processing.
[0036] When obtaining the first image for subsequent processing based on the video frame collected by the endoscope, reference can be made to Figure 2 , Figure 2 which is a schematic flowchart of obtaining the first image provided by an embodiment of the present application. Among them, this step includes steps 201 to 203.
[0037] Step 201, obtain a video frame collected by the endoscope, and perform marker recognition on the video frame to determine whether the video frame contains a marker;
[0038] Step 202, when it is determined that the video frame contains a marker, identify the invalid area of the video frame;
[0039] Step 203, perform image cropping on the video frame according to the identified invalid area to obtain the processed first image.
[0040] Specifically, for the video frames collected by the endoscope, not all of them can be used for subsequent processing, that is, not all video frames contain markers such as lesions. At the same time, for the obtained video frames containing markers such as lesions, they will also contain other invalid information.
[0041] Therefore, after obtaining the video frames collected by the endoscope, perform marker recognition on the fragmented video frames to determine whether the video frames contain markers. Then, in the case where it is determined that the video frames contain markers, further process the video frames to remove the invalid information in the video frames and obtain the first image for subsequent processing.
[0042] Exemplarily, the acquisition of video frames is carried out in real time, and the same processing is performed for each video frame during analysis, including: determining whether it contains markers and cropping the video frames containing markers.
[0043] When determining whether a video frame contains markers, it can be determined by using a pre-trained related model. For example, a lesion segmentation neural network is pre-trained to identify lesions in the video frame. The constructed training image set is used to train the constructed lesion segmentation neural network. The constructed training image set can be an image set manually labeled with lesions, an image set labeled for the same type of lesion, or an image set labeled for multiple types of lesions.
[0044] When performing cropping processing on the video frame, just remove the invalid information contained in it. Specifically, it includes: cropping the video frame, removing the identified invalid area from the video frame to obtain the cropped video frame; performing image scaling processing on the cropped video frame, and obtaining the first image after completing the scaling processing.
[0045] That is, when performing cropping processing on a video frame containing markers to obtain the first image, first identify the invalid information in the video frame, and at the same time determine the area where the invalid information is located. Then, crop and remove the area where the invalid information is located. After completing the cropping of the video frame, perform image scaling processing on the cropped video frame, and then obtain the first image after completing the scaling processing.
[0046] When cropping a video frame containing markers, the specific cropping process is as Figure 3 , in Figure 3 the above video frame in the figure, it contains relevant information such as device information and parameters, which are pre-set. And these information are invalid when identifying the lesions contained in the image. Therefore, determine an invalid area for the invalid information, and then crop the invalid area to obtain, for example,Figure 3 The video frame shown in the middle and bottom figures.
[0047] In addition, after completing the cropping of the video frame, in order to ensure the unity of image processing, at this time, the cropped video frame can also be subjected to image scaling processing, such as scaling the cropped video frame to a size of 512×512.
[0048] Furthermore, when performing the cropping process on the video frame, the area interpolation method can be used for processing. As Figure 4 shown, when the picture is scaled down, the pixel point (x′, y′) of the scaled-down picture corresponds to the upper left corner on the original picture as (x′×scale x , y′×scale y ), and the lower right corner is ((x′ + 1)×scale x - 1, (y′ + 1)×scale y - 1) such a pixel area Area. Among them, scale x , scale y is the multiple of the width and height of the original picture divided by the width and height of the scaled-down picture. When it cannot be divided evenly, this multiple is a decimal. The pixel value of the pixel point (x′, y′) is the average value of the pixel values of all points in the pixel area included in the original picture.
[0049] When the scaling factor is not an integer, as Figure 5 shown, only a part of the edge pixels may be included in the pixel area. At this time, the weight of the pixels completely included is 1, and the partially included pixels are weighted according to the included ratio. Thus, the formula expression of the area interpolation method is as follows:
[0050]
[0051] Among them, scale_x and scale_y are the multiples of the width and height of the original picture divided by the width and height of the scaled-down picture, Weight(x, y) is the ratio of the pixel (x, y) on the original picture included in the pixel area, and Area is the area of the pixel area.
[0052] Step 102: Extract features from the first image to obtain the marker size information and reference object size information corresponding to the first image.
[0053] In one embodiment, after obtaining the first image, feature extraction is performed on the first image to obtain the size information of the marker corresponding to the first image and the size information of the reference object. It should be noted that for the endoscope used for video frame acquisition, a corresponding reference object is provided on the endoscope. When video frames are acquired, video frames containing the reference object are obtained. At the same time, when processing the video frames, the reference object contained in the video frames is not removed. In various embodiments, the reference object may be a transparent cap provided on the endoscope. Among them, the medical name of the transparent cap is "medical transparent mucosal suction sleeve", which is used in combination with the endoscope and is assembled at the front end of the endoscope lens for maintaining an appropriate endoscopic field of view when observing the lumen wall. Since the relative positional relationship between the transparent cap and the endoscope is fixed and known, when the transparent cap is used as the reference object, the true size information of markers such as lesions can be more accurately compared and calculated, and no other additional components are required as the reference object.
[0054] Specifically, when extracting the feature information of the first image to obtain the marker size information and the reference object size information, it includes: detecting the reference object in the first image to obtain a first prediction probability map, and based on the first prediction probability map, obtaining the reference object size information of the first image; detecting the marker in the first image to obtain a second prediction probability map, and based on the second prediction probability map, obtaining the marker size information of the first image.
[0055] Exemplarily, when processing the first image to obtain the marker size information and the reference object size information, it is processed based on the respective set methods.
[0056] Taking the obtaining of the reference object size information as an example, by detecting the reference object in the first image, a first prediction probability map for the reference object is obtained, and then through the analysis and processing of the first probability prediction map, the reference object size information of the first image is obtained. Among them, when detecting the reference object in the first image, it can be obtained by processing using a pre-trained related model, such as a transparent cap segmentation neural network. Similarly, it is trained using a pre-annotated image set, and then when detecting the transparent cap, mask processing is performed on the pixel points where the detected reference object is located to obtain the first prediction probability map.
[0057] Next, the reference object size information is obtained by analyzing and processing the first prediction probability map. Specifically, referring to Figure 6 , Figure 6 is a schematic flowchart of the steps for obtaining the reference object size information provided by the embodiments of the present application, where this step includes steps 601 to 603.
[0058] Step 601, perform binarization processing on the first prediction probability map, and obtain a second image after completing the binarization processing;
[0059] Step 602: Perform pixel restoration processing on the second image to obtain a third image;
[0060] Step 603: Identify and extract features of the reference object included in the third image to obtain the reference object size information of the reference object included in the first image.
[0061] Exemplarily, after obtaining the first prediction probability map, through the processing of the first probability prediction map, the reference object size information of the reference object included in the first image can be obtained. When performing the processing, first perform binarization processing on the first probability prediction map to obtain a second image, then perform pixel restoration processing on the second image to obtain a third image, and finally identify and extract features of the reference object included in the third image to obtain the reference object size information of the reference object included in the first image.
[0062] In the actual processing process, taking the reference object as a transparent cap as an example, after using the pre-trained transparent cap segmentation neural network to detect the first image to obtain the first prediction probability map, perform binarization processing on the first prediction probability map through the set threshold Thr to obtain a grayscale image G a , where Thr can be set to 0.5, that is:
[0063]
[0064] Then, perform pixel value restoration processing on G a :
[0065] G a = G a × 255;
[0066] Next, the opening circle and the pixel diameter of the opening of the transparent cap can be identified through the Hough transform circle detection. Then, perform recognition processing on the pixel-restored image to recognize all the digital scales in the image. The actual circumference of the circle of the transparent cap is the maximum value among all the scale values, that is
[0067] C = Max(V);
[0068] That is, the reference object size information C of the reference object included in the first image is obtained.
[0069] Taking the acquisition of the marker size information as an example, by detecting the marker in the first image, a second predicted probability map for the marker is obtained. Herein, the marker may be a lesion, and then through the analysis and processing of the second probability prediction map, the marker size information of the first image is obtained. Among them, when detecting the marker in the first image, it can be obtained by processing with a pre-trained related model, such as the aforementioned lesion segmentation neural network. When performing lesion detection, mask processing is performed on the pixel points where the detected lesions are located to obtain the second predicted probability map.
[0070] Next, by analyzing and processing the second predicted probability map, the marker size information is obtained. Specifically, referring to Figure 7 , Figure 7 is a schematic flowchart of the steps for obtaining the marker size information provided by an embodiment of the present application. Among them, this step includes steps 701 to 703.
[0071] Step 701: Perform binarization processing on the second predicted probability map, and obtain a fourth image after the binarization processing is completed;
[0072] Step 702: Perform pixel reduction processing on the fourth image to obtain a fifth image;
[0073] Step 703: Identify and extract features of the markers included in the fifth image to obtain a marker set included in the fifth image, and select the size information of the target reference object in the marker set as the marker size information included in the first image, wherein the corresponding relationship between the markers and the size information is recorded in the marker set.
[0074] Exemplarily, after obtaining the second predicted probability map, through the processing of the second probability prediction map, the marker size information of the markers included in the second image can be obtained. When performing the processing, first, binarization processing is performed on the second probability prediction map to obtain a fourth image, then pixel reduction processing is performed on the fourth image to obtain a fifth image, and finally, the markers included in the fifth image are identified and feature-extracted to obtain a marker set included in the fifth image, and then the marker currently being identified and processed is determined in the marker set to obtain the required size information at this time as the marker size information included in the first image.
[0075] In the actual processing process, taking the marker as a lesion as an example, after using the pre-trained lesion segmentation neural network to detect the second image to obtain the second probability prediction map, binarization processing is performed on the second predicted probability map through a set threshold Thr to obtain a grayscale image G c , where Thr can be set to 0.48, that is:
[0076]
[0077] Then, perform pixel value restoration on G c :
[0078] G c = G c × 255;
[0079] Next, perform connected component detection on G c to obtain a set of connected components C c , and use C ci to represent the set of points in each recognized connected component. Perform outlier detection on the points in each C ci , such as the LOP (Local Outlier Factor) algorithm. The set of connected components after removing outliers is represented by C cr .
[0080] By calculating the area of each connected component in C cr and sorting them in descending order, obtain a set of connected component indices I cr .
[0081] Extract the required number n of lesion-connected components C' c
[0082] C' c = C cr [I cri ;
[0083] where i ∈ [0, n].
[0084] Perform minimum bounding rectangle calculation on each connected component in C' c to obtain a set of rectangles S, that is, obtain a set of markers, and further obtain the pixel length and width L' a , L' b of the required n lesions. For example, when it is necessary to confirm and identify the largest lesion, n = 1 can be taken. Then, when obtaining the marker size information, the pixel length and width of the largest lesion will be obtained.
[0085] Step 103: Analyze the video frame to obtain the distance value of the target marker corresponding to the endoscope and marker size information.
[0086] In one embodiment, while processing the first image, the video frames collected by the endoscope can also be analyzed to obtain the distance value of the target marker corresponding to the endoscope and the marker size information. Specifically, when analyzing the video frames to obtain the distance value between the endoscope and the target marker, it can be processed using a pre-trained relevant model, such as a pre-trained depth estimation neural network. The model is trained and optimized using a set of annotated relevant images. Then, after obtaining the trained relevant model, the video frames are analyzed to determine the distance value between the endoscope and the target marker.
[0087] Step 104: Obtain the true size information of the target marker based on the distance value, the marker size information, and the reference object size information.
[0088] In one embodiment, after obtaining the reference object size information and the marker size information included in the first image, the distance value between the endoscope and the target marker obtained by analyzing the video frames is combined for corresponding transformation and calculation to obtain the true size information of the currently recognized target marker.
[0089] Exemplarily, when calculating the true size information of the target marker, it includes: determining the first plane to which the marker corresponding to the marker size information belongs, and determining the second plane to which the reference object corresponding to the reference object size information belongs; aligning the marker size information and the reference object size information according to the first plane and the second plane; obtaining the position information of the target reference object corresponding to the reference object size information on the endoscope, and calculating the true size information of the target marker based on the aligned marker size information, the reference object size information, the distance value, and the position information.
[0090] In the actual processing, since the planes where the marker and the reference object are located may be different, when calculating the true size information of the marker through comparison, the marker and the reference object need to be transformed to the same plane and then calculated to obtain the true size information of the marker.
[0091] Taking the above-described specific example as an example, after obtaining the distance L between the endoscope and the marker d Since the position of the marker on the endoscope is fixed, the height of the transparent cap (marker) is a fixed value L c ,
[0092] Then, the pixel length and width of the lesion are converted to the same plane as the opening of the transparent cap to obtain the converted pixel length and width L′′ of the lesion (marker) a , L′′ b .
[0093]
[0094]
[0095] Next, according to the actual circumference C of the opening circle of the transparent cap, the actual length and width L of the lesion are calculated. a , L b .
[0096]
[0097]
[0098] That is, the true size information of the marker is obtained as L a and L b .
[0099] In summary, the present application discloses a marker recognition method. In the process of recognizing and processing markers such as lesions, video frames collected by the inner diameter are obtained, and the collected video frames are screened and processed to obtain a first image for marker recognition. Then, feature information of the first image is extracted to obtain the marker size information and reference object size information corresponding to the first image. At the same time, by analyzing the video frames, the distance relationship between the currently recognized marker and the inner diameter, such as the distance value, is determined. Furthermore, according to the obtained marker size information, reference object size information, and distance relationship, the true size information of the target marker included in the first image is determined. It realizes that in the recognition and processing process, a reference object is added to recognize the marker and the reference object included in the collected video frames, and then uses the set information such as the position of the reference object to obtain the true size information of the marker, improving the accuracy of recognizing markers such as lesions.
[0100] According to the method described in the above embodiments, this embodiment will be further described from the perspective of the marker recognition device. The marker recognition device can be specifically implemented as an independent entity or integrated in an electronic device, such as a terminal. The terminal can include a mobile phone, a tablet computer, etc.
[0101] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of the marker recognition device provided by the embodiment of the present application. As Figure 8 shown, the marker recognition device 800 provided by the embodiment of the present application includes:
[0102] A video acquisition module 801, configured to obtain video frames collected by the endoscope and process the video frames to obtain a first image;
[0103] A feature extraction module 802, configured to extract features from the first image to obtain the marker size information and reference object size information corresponding to the first image;
[0104] The video analysis module 803 is configured to analyze video frames to obtain the distance value of the target marker corresponding to the endoscopic and marker size information;
[0105] The result processing module 804 is configured to obtain the true size information of the target marker according to the distance value, the marker size information, and the reference object size information.
[0106] In one embodiment, the video acquisition module 801 is further configured to:
[0107] Acquire the video frames collected by the endoscope, perform marker recognition on the video frames, and determine whether the video frames contain markers;
[0108] When it is determined that the video frames contain markers, identify the invalid regions of the video frames;
[0109] Perform image cropping on the video frames according to the identified invalid regions to obtain the processed first image.
[0110] In one embodiment, the video acquisition module 801 is further configured to:
[0111] Crop the video frames to remove the identified invalid regions from the video frames to obtain the cropped video frames;
[0112] Perform image scaling processing on the cropped video frames, and obtain the first image after the scaling processing is completed.
[0113] In one embodiment, the feature extraction module 802 is further configured to:
[0114] Perform reference object detection on the first image to obtain the first prediction probability map, and obtain the reference object size information of the first image according to the first prediction probability map;
[0115] Perform marker detection on the first image to obtain the second prediction probability map, and obtain the marker size information of the first image according to the second prediction probability map.
[0116] In one embodiment, the feature extraction module 802 is further configured to:
[0117] Perform binarization processing on the first prediction probability map, and obtain the second image after the binarization processing is completed;
[0118] Perform pixel reduction processing on the second image to obtain the third image;
[0119] Identify and extract the features of the reference object included in the third image to obtain the reference object size information of the reference object included in the first image.
[0120] In one embodiment, the feature extraction module 802 is further configured to:
[0121] Binarize the second predicted probability map, and obtain a fourth image after the binarization process is completed;
[0122] Perform pixel reduction processing on the fourth image to obtain a fifth image;
[0123] Identify and extract features of the markers included in the fifth image to obtain a set of markers included in the fifth image, and select the size information of the target reference object in the set of markers as the size information of the markers included in the first image, where the corresponding relationship between the markers and the size information is recorded in the set of markers.
[0124] In one embodiment, the result processing module 804 is further configured to:
[0125] Determine the first plane to which the marker corresponding to the marker size information belongs, and determine the second plane to which the reference object corresponding to the reference object size information belongs;
[0126] Align the marker size information and the reference object size information according to the first plane and the second plane;
[0127] Obtain the position information of the target reference object corresponding to the reference object size information on the endoscope, and calculate the true size information of the target marker based on the aligned marker size information and reference object size information based on the distance value and the position information.
[0128] In addition, please refer to Figure 9 , Figure 9 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device may be a mobile terminal such as a smart phone, a tablet computer, etc. As Figure 9 shown, the electronic device 900 includes a processor 901 and a memory 902. Among them, the processor 901 is electrically connected to the memory 902.
[0129] The processor 901 is the control center of the electronic device 900, connects various parts of the entire electronic device through various interfaces and lines, executes various functions of the electronic device 900 and processes data by running or loading application programs stored in the memory 902, and calling data stored in the memory 902, thereby monitoring the electronic device 900 as a whole.
[0130] In this embodiment, the processor 901 in the electronic device 900 will load the instructions corresponding to the processes of one or more application programs into the memory 902 according to the following steps, and the processor 901 will run the application programs stored in the memory 902 to implement any step in the marker recognition method provided in the above embodiment.
[0131] The electronic device 900 may implement the steps in any of the embodiments of the marker recognition method provided in the embodiments of the present application. Therefore, it can achieve the beneficial effects that any of the marker recognition methods provided in the embodiments of the present application can achieve. For details, please refer to the previous embodiments and will not be elaborated here.
[0132] Please refer to Figure 10 , Figure 10 which is another structural schematic diagram of the electronic device provided in the embodiments of the present application. As Figure 10 shown, Figure 10 it shows the specific structural block diagram of the electronic device provided in the embodiments of the present application. The electronic device 1000 may be a mobile terminal such as a smart phone or a laptop computer.
[0133] The RF circuit 1010 is used to receive and transmit electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices. The RF circuit 1010 may include various existing circuit elements for performing these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The RF circuit 1010 can communicate with various networks such as the Internet, enterprise intranets, wireless networks or communicate with other devices through a wireless network. The above-mentioned wireless network may include a cellular phone network, a wireless local area network or a metropolitan area network. The above-mentioned wireless network can use various communication standards, protocols and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as Institute of Electrical and Electronics Engineers standards IEEE 802.11a, IEEE 802.11b, IEEE 802.11g and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging and short messages, and any other suitable communication protocols, and may even include those protocols that have not yet been developed currently.
[0134] The memory 1020 can be used to store software programs and modules, such as the program instructions / modules corresponding to the marker recognition method in the above embodiments. The processor 1080 executes various functional applications and the marker recognition method by running the software programs and modules stored in the memory 1020.
[0135] The memory 1020 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some examples, the memory 1020 may further include a memory remotely located relative to the processor 1080, and these remote memories may be connected to the electronic device 1000 through a network. Examples of the above network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0136] The input unit 1030 can be used to receive uploaded digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls. Specifically, the input unit 1030 may include a touch-sensitive surface 1031 and other input devices 1032. The touch-sensitive surface 1031, also known as a touch display screen or touchpad, can collect touch operations of the user on or near it (such as operations of the user using a finger, stylus, or any suitable object or accessory on or near the touch-sensitive surface 1031), and drive the corresponding connection device according to a preset program. Optionally, the touch-sensitive surface 1031 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1080, and can receive and execute commands sent by the processor 1080. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch-sensitive surface 1031. In addition to the touch-sensitive surface 1031, the input unit 1030 may further include other input devices 1032. Specifically, the other input devices 1032 may include but are not limited to one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), trackballs, mice, joysticks, etc.
[0137] The display unit 1040 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device 1000. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. The display unit 1040 may include a display panel 1041. Optionally, the display panel 1041 can be configured in the form of an LCD (Liquid Crystal Display), an OLED (Organic Light-Emitting Diode), etc. Further, the touch-sensitive surface 1031 can cover the display panel 1041. When the touch-sensitive surface 1031 detects a touch operation on or near it, it is transmitted to the processor 1080 to determine the type of touch event. Subsequently, the processor 1080 provides a corresponding visual output on the display panel 1041 according to the type of touch event. Although in the figure, the touch-sensitive surface 1031 and the display panel 1041 are implemented as two independent components to achieve the input and output functions, in some embodiments, the touch-sensitive surface 1031 and the display panel 1041 can be integrated to achieve the input and output functions.
[0138] The electronic device 1000 may further include at least one sensor 1050, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1041 according to the brightness of the ambient light, and the proximity sensor can generate an interruption when the flip cover is closed or opened. As a kind of motion sensor, the gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used in applications for identifying the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, tapping), etc. As for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the electronic device 1000 can also be configured with, they will not be elaborated here.
[0139] The audio circuit 1060, the speaker 1061, and the microphone 1062 can provide an audio interface between the user and the electronic device 1000. The audio circuit 1060 can transmit the electrical signal converted from the received audio data to the speaker 1061, and the speaker 1061 converts it into a sound signal for output. On the other hand, the microphone 1062 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1060 and then converted into audio data. After the audio data is output to the processor 1080 for processing, it is sent to another terminal, for example, via the RF circuit 1010, or the audio data is output to the memory 1020 for further processing. The audio circuit 1060 may also include an earphone jack to provide communication between the external earphone and the electronic device 1000.
[0140] The electronic device 1000 can help users receive requests, send information, etc. through a transmission module 1070 (such as a Wi-Fi module), providing users with wireless broadband Internet access. Although the transmission module 1070 is shown in the figure, it can be understood that it does not belong to an essential component of the electronic device 1000 and can be omitted entirely within the scope of not changing the essence of the invention as needed.
[0141] The processor 1080 is the control center of the electronic device 1000, connecting various parts of the entire mobile phone using various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 1020, and by invoking data stored in the memory 1020, it performs various functions of the electronic device 1000 and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 1080 may include one or more processing cores; in some embodiments, the processor 1080 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1080.
[0142] The electronic device 1000 also includes a power supply 1090 (such as a battery) for powering each component. In some embodiments, the power supply can be logically connected to the processor 1080 through a power management system, thereby realizing functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 1090 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator, etc.
[0143] Although not shown, the electronic device 1000 also includes a camera (such as a front camera, a rear camera), a Bluetooth module, etc., which will not be elaborated here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the mobile terminal also includes a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors to implement any step in the marker recognition method provided in the above embodiments.
[0144] During specific implementation, the above-mentioned each module can be implemented as an independent entity, or can be combined arbitrarily to be implemented as the same or several entities. For the specific implementation of the above-mentioned each module, reference can be made to the method embodiments described above, which will not be elaborated here.
[0145] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. To this end, an embodiment of the present application provides a storage medium in which multiple instructions are stored. When the instructions can be executed by a processor, any step in the marker recognition method provided in the above embodiments can be implemented.
[0146] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0147] Since the instructions stored in the storage medium can execute the steps in any embodiment of the marker recognition method provided in the embodiments of the present application, the beneficial effects achievable by any marker recognition method provided in the embodiments of the present application can be achieved. For details, see the previous embodiments and will not be repeated here.
[0148] The above has introduced in detail a marker recognition method, device, electronic device and storage medium provided by the embodiments of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application. Moreover, for those of ordinary skill in the technical field, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present application.
Claims
1. A marker identification method, characterized in that: include: Acquire a video frame captured by the endoscope, and process the video frame to obtain a first image; Performing reference object detection on the first image to obtain a first prediction probability map, and obtaining reference object size information of the first image based on the first prediction probability map, and performing marker detection on the first image to obtain a second prediction probability map, and obtaining marker size information of the first image based on the second prediction probability map, wherein the reference object is set on the endoscope; Analyzing the video frame to obtain a distance value between the endoscope and a target marker corresponding to the marker size information; Obtaining the real size information of the target marker according to the distance value, the marker size information and the reference object size information; The step of performing marker detection on the first image to obtain a second prediction probability map, and obtaining marker size information of the first image according to the second prediction probability map includes: Binarizing the second prediction probability map, and obtaining a fourth image after the binarization is completed; Performing pixel restoration processing on the fourth image to obtain a fifth image; Identify and extract features of the markers contained in the fifth image to obtain a marker set contained in the fifth image, and select size information of a target reference object from the marker set as the marker size information contained in the first image, wherein the marker set records the correspondence between the marker and the size information; Furthermore, the identifying and feature extracting the markers contained in the fifth image to obtain a marker set contained in the fifth image, and selecting size information of a target reference object from the marker set as the marker size information contained in the first image, includes: Performing a connected domain detection on the fifth image to obtain a plurality of connected domains contained in the fifth image to form a connected domain set, wherein one connected domain formed by performing the connected domain detection corresponds to one marker; Calculate the connected domain area of each connected domain in the connected domain set, and sort the connected domain areas from large to small to obtain a connected domain index set, wherein the connected domain area is the minimum rectangular area of the connected domain; When it is determined to select the maximum marker, the maximum connected domain area is used as the marker size information in the connected domain index set.
2. The method according to claim 1, characterized in that The step of acquiring the video frame collected by the endoscope and processing the video frame to obtain the first image includes: Acquire a video frame captured by the endoscope, and perform landmark recognition on the video frame to determine whether the video frame contains a landmark; When it is determined that the video frame contains a marker, identifying an invalid area of the video frame; The video frame is cropped according to the identified invalid area to obtain a processed first image.
3. The method according to claim 2, characterized in that The step of cropping the video frame according to the identified invalid area to obtain a processed first image includes: Cropping the video frame, removing the identified invalid area from the video frame, and obtaining a cropped video frame; The cropped video frame is subjected to image scaling processing, and a first image is obtained after the scaling processing is completed.
4. The method according to claim 1, characterized in that The obtaining the reference object size information of the first image according to the first prediction probability map includes: Binarizing the first prediction probability map, and obtaining a second image after the binarization is completed; Performing pixel restoration processing on the second image to obtain a third image; The reference object contained in the third image is identified and features are extracted to obtain reference object size information of the reference object contained in the first image.
5. The method according to claim 1, characterized in that The obtaining the real size information of the target marker according to the distance value, the marker size information and the reference object size information includes: Determine a first plane to which the marker corresponding to the marker size information belongs, and determine a second plane to which the reference object corresponding to the reference object size information belongs; Performing plane alignment on the marker size information and the reference size information according to the first plane and the second plane; The position information of the target reference object on the endoscope corresponding to the reference object size information is obtained, and the real size information of the target marker is calculated based on the distance value and the position information according to the aligned marker size information and the reference object size information.
6. A marker recognition device, characterized in that: include: A video acquisition module, used to acquire video frames acquired by the endoscope, and process the video frames to obtain a first image; a feature extraction module, configured to perform reference object detection on the first image to obtain a first prediction probability map, and obtain reference object size information of the first image based on the first prediction probability map, and perform marker detection on the first image to obtain a second prediction probability map, and obtain marker size information of the first image based on the second prediction probability map, wherein the reference object is set on the endoscope; A video analysis module, used for analyzing the video frame to obtain a distance value between the endoscope and a target marker corresponding to the marker size information; A result processing module, used for obtaining the real size information of the target marker according to the distance value, the marker size information and the reference object size information; Wherein, the feature extraction module is also used for: Binarizing the second prediction probability map, and obtaining a fourth image after the binarization is completed; Performing pixel restoration processing on the fourth image to obtain a fifth image; Identify and extract features of the markers contained in the fifth image to obtain a marker set contained in the fifth image, and select size information of a target reference object from the marker set as the marker size information contained in the first image, wherein the marker set records the correspondence between the marker and the size information; Furthermore, the feature extraction module is further used for: Performing a connected domain detection on the fifth image to obtain a plurality of connected domains contained in the fifth image to form a connected domain set, wherein one connected domain formed by performing the connected domain detection corresponds to one marker; Calculate the connected domain area of each connected domain in the connected domain set, and sort the connected domain areas from large to small to obtain a connected domain index set, wherein the connected domain area is the minimum rectangular area of the connected domain; When it is determined to select the maximum marker, the maximum connected domain area is used as the marker size information in the connected domain index set.
7. An electronic device, characterized in that: The electronic device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor implements the steps in the method according to any one of claims 1 to 6 when executing the computer program.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Fruit and vegetable size identification method and device, electronic equipment and computer readable medium
CN112257506A
Method and device for adding AR explanation based on object positioning
CN115797602A