Image processing method and device, storage medium and electronic equipment
By generating an endoscopic image of the target object in the center of the field of view through an image supplementation and generation network, the problem of inaccurate target object location identification in endoscopic examination is solved, and the accuracy of location identification is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- RENMIN HOSPITAL OF WUHAN UNIVERSITY (HUBEI GENERAL HOSPITAL)
- Filing Date
- 2023-02-22
- Publication Date
- 2026-05-01
AI Technical Summary
The accuracy of target object location identification is low during endoscopic examinations, especially under the heavy workload of physicians. Target objects are often displayed outside the center of the field of vision in the image, leading to inaccurate location judgment.
By using image processing methods, an endoscopic image of the target object located in the center of the field of view is generated using an image supplementation network and a generative network, and then a location recognition network is used to obtain accurate target location information.
It improves the accuracy of target object location recognition in endoscopic images, ensuring that the target object is clearly displayed in the center of the field of view, making it easier for physicians to accurately determine its location.
Smart Images

Figure CN116309372B_ABST
Abstract
Description
Image processing methods, apparatuses, storage media and electronic devices Technical Field
[0001] This application relates to the field of image processing technology, and in particular to an image processing method, apparatus, storage medium and electronic device. Background Technology
[0002] With the continuous development of medical technology, using endoscopy to identify abnormalities in the body has become a relatively mature and reliable medical method.
[0003] Endoscopic examinations not only require physicians to possess high levels of medical expertise and clinical experience, but also place high demands on their physical and mental well-being. If physicians use endoscopes to examine abnormalities under heavy workloads, the quality of the examination may be reduced. For example, if the lens is too close to the tissue, the tissue may not be fully displayed in the endoscopic image. Or, if the target object (e.g., an abnormal area) is located at the edge of the tissue, a poor shooting angle may cause the target object to appear at the edge of the endoscopic image (i.e., not in the center of the endoscopic field of view). Physicians may not be able to accurately determine the specific location of the target object within the tissue based on the endoscopic image, resulting in a low accuracy rate in identifying the target object's location. Summary of the Invention
[0004] This application provides an image processing method, apparatus, storage medium, and electronic device to alleviate the technical problem of low accuracy in identifying the location of a target object.
[0005] To address the aforementioned technical problems, this application provides the following technical solution:
[0006] This application provides an image processing method, including:
[0007] Acquire an endoscope image to be processed; wherein the target object in the endoscope image to be processed is located in a non-preset area of the endoscope image to be processed;
[0008] The endoscope image to be processed is input into an image supplementation network to obtain a reference endoscope image;
[0009] The endoscope image to be processed and the reference endoscope image are input into an image generation network to generate a target endoscope image; wherein the target object is located within a preset area of the target endoscope image;
[0010] The endoscope image to be processed and the target endoscope image are input into a location recognition network to obtain the target location information of the target object.
[0011] The step of acquiring the endoscopic image to be processed includes:
[0012] Multiple initial endoscopic images are input into a target object recognition network to obtain intermediate endoscopic images; wherein, the intermediate endoscopic images include the initial endoscopic images containing the target object;
[0013] If the target object in the intermediate endoscopic image is located within the non-preset area of the intermediate endoscopic image, the intermediate endoscopic image is used as the endoscopic image to be processed.
[0014] The step of inputting multiple initial endoscopic images into a target object recognition network to obtain intermediate endoscopic images includes:
[0015] Multiple initial endoscopic images are input into a target object recognition network to extract features from each initial endoscopic image through the target object recognition network, thereby obtaining the features of the initial endoscopic images.
[0016] If the initial endoscopic image features are target object features, then the initial endoscopic image is determined to be the intermediate endoscopic image.
[0017] The step of inputting the endoscope image to be processed into an image supplementation network to obtain a reference endoscope image includes:
[0018] The endoscopic image to be processed is input into the image supplementation network to extract tissue features from the endoscopic image to be processed, and the reference endoscopic image is determined based on the tissue features.
[0019] The step of inputting the endoscopic image to be processed into the image supplementation network, extracting tissue features from the endoscopic image to be processed through the image supplementation network, and determining the reference endoscopic image based on the tissue features includes:
[0020] The endoscope image to be processed is input into the image supplementation network to extract tissue features in the endoscope image to be processed. If the tissue feature is a reference tissue feature, the endoscope image corresponding to the reference tissue feature is determined.
[0021] The endoscopic image corresponding to the features of the reference tissue site is used as the reference endoscopic image.
[0022] The step of inputting the endoscope image to be processed and the reference endoscope image into an image generation network to generate a target endoscope image includes:
[0023] In the endoscopic image to be processed, determine the region image of the area where the target object is located, and extract the tissue mucosal features of the region image;
[0024] The endoscope image to be processed, the reference endoscope image, and the tissue mucosal features are input into the image generation network to determine the stitching position of the target object in the reference endoscope image. The endoscope image to be processed and the reference endoscope image are stitched together according to the stitching position to obtain a stitched image. A tissue mucosal object is generated in the stitched image according to the tissue mucosal features.
[0025] The stitched image containing the tissue mucosa object is used as the target endoscopic image.
[0026] The location recognition network includes a first location recognition network and a second location recognition network. The step of inputting the endoscope image to be processed and the target endoscope image into the location recognition network to obtain the target location information of the target object includes:
[0027] The endoscope image to be processed is input into the first location recognition network to obtain the first location information of the target object, and the target endoscope image is input into the second location recognition network to obtain the second location information of the target object;
[0028] The target location information of the target object is determined based on the first location information and the second location information.
[0029] This application also provides an image processing apparatus, including:
[0030] An acquisition module is used to acquire an endoscope image to be processed; wherein the target object in the endoscope image to be processed is located in a non-preset area of the endoscope image to be processed;
[0031] The reference endoscope image determination module is used to input the endoscope image to be processed into the image supplementation network to obtain the reference endoscope image;
[0032] A generation module is used to input the endoscope image to be processed and the reference endoscope image into an image generation network to generate a target endoscope image; wherein the target object is located within a preset area of the target endoscope image;
[0033] The target location information determination module is used to input the endoscope image to be processed and the target endoscope image into the location recognition network to obtain the target location information of the target object.
[0034] This application also provides a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to perform steps in any of the above-described image processing methods.
[0035] This application also provides an electronic device, including a processor and a memory, wherein the processor is electrically connected to the memory, the memory is used to store instructions and data, and the processor is used to execute the steps in any of the above-described image processing methods.
[0036] This application provides an image processing method, apparatus, storage medium, and electronic device. First, an endoscope image to be processed is acquired, wherein the target object in the endoscope image is located outside a preset region of the image. Then, the endoscope image to be processed is input into an image supplementation network to obtain a reference endoscope image. Next, the endoscope image to be processed and the reference endoscope image are input into an image generation network to generate a target endoscope image, wherein the target object is located within a preset region of the target endoscope image. Finally, the endoscope image to be processed and the target endoscope image are input into a location recognition network to obtain the target location information of the target object. Based on the endoscope image to be processed, a target endoscope image is automatically generated where the target object is located within a preset region of the field of view (e.g., the center of the field of view). The location information of the target object is accurately obtained based on its position in the endoscope image to be processed and the target endoscope image, thereby effectively improving the accuracy of target object location recognition and alleviating the current technical problem of low target object location recognition accuracy. Attached Figure Description
[0037] The technical solution and other beneficial effects of this application will become apparent from the following detailed description of specific embodiments in conjunction with the accompanying drawings.
[0038] Figure 1 is a flowchart illustrating the image processing method provided in an embodiment of this application.
[0039] Figure 2 is a schematic diagram of a scene of the image processing method provided in an embodiment of this application.
[0040] Figure 3 is a schematic diagram of another scenario of the image processing method provided in the embodiments of this application.
[0041] Figure 4 is another scenario diagram of the image processing method provided in the embodiments of this application.
[0042] Figure 5 is a schematic diagram of the structure of the image processing device provided in the embodiment of this application.
[0043] Figure 6 is a schematic diagram of the structure of the electronic device provided in an embodiment of this application.
[0044] Figure 7 is another structural schematic diagram of the electronic device provided in an embodiment of this application. Detailed Implementation
[0045] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0046] This application provides an image processing method, apparatus, storage medium, and electronic device.
[0047] As shown in Figure 1, Figure 1 is a schematic flowchart of the image processing method provided in an embodiment of this application. The specific process can be as follows:
[0048] S101. Obtain the endoscope image to be processed, wherein the target object in the endoscope image to be processed is located in a non-preset area of the endoscope image to be processed.
[0049] The endoscopic images to be processed include endoscopic images of the target object located outside the preset area of the image. The target object includes abnormal areas in the endoscopic images. Optionally, the abnormal area can be an area in the endoscopic image that is different from the surrounding area. For example, if the color of the first area in the endoscopic image is different from the color of the surrounding area, then the first area is considered an abnormal area. In this scenario, foreign objects can be detected. For example, if a foreign object protrudes into the esophagus, since the color of the foreign object is different from the color of the esophagus, the location of the foreign object can be located by comparing the colors of the areas. Another example is the duodenum. If a second area in the duodenal region shown in the endoscopic image is in a convex state, and the degree of convexity of the second area is different from that of the surrounding area, then the second area is considered an abnormal area.
[0050] Specifically, in practical applications, endoscopes are typically used to examine multiple parts of the upper digestive tract (e.g., esophagus, cardia, greater curvature of the gastric antrum, lesser curvature of the gastric antrum, anterior wall of the gastric antrum, posterior wall of the gastric antrum, lower greater curvature of the gastric body, lower lesser curvature of the gastric body, anterior wall of the lower gastric body, posterior wall of the lower gastric body, upper-middle greater curvature of the gastric body, upper-middle lesser curvature of the gastric body, upper-middle anterior wall of the gastric body, upper-middle posterior wall of the gastric body, inverted lesser curvature of the gastric angle, anterior wall of the inverted gastric angle, posterior wall of the inverted gastric angle, upper-middle lesser curvature of the inverted gastric body, upper-middle anterior wall of the inverted gastric body, upper-middle posterior wall of the inverted gastric body, greater curvature of the gastric fundus, lesser curvature of the gastric fundus, anterior wall of the inverted gastric fundus, posterior wall of the inverted gastric fundus, duodenal bulb, descending duodenum) for examination.
[0051] However, if the target object is located at the edge of the upper gastrointestinal tract, or when physicians perform endoscopic examinations under high workload, the acquired endoscopic images often exhibit incomplete display of tissue areas, with the target object located within the tissue area appearing in the edge region (i.e., a non-central region) of the endoscopic image. For example, as shown in Figure 2, because the target object 2001 is located at the edge of the gastric antrum structure 2002, and because the lens is too close to the gastric antrum structure 2002 during endoscopic image acquisition, the gastric antrum structure 2002 is not fully displayed in the acquired endoscopic image, and the target object 2001 is located in a non-central region of the endoscopic image. Optionally, in this embodiment, the non-preset region is the non-central region of the field of view, and the endoscopic image to be processed is the endoscopic image exhibiting the aforementioned incomplete display phenomenon.
[0052] Furthermore, step S101 specifically includes:
[0053] Multiple initial endoscopic images are input into a target object recognition network to obtain intermediate endoscopic images, wherein the intermediate endoscopic images include initial endoscopic images with the target object.
[0054] If the target object in the intermediate endoscope image is located in a non-preset area of the intermediate endoscope image, the intermediate endoscope image is used as the endoscope image to be processed.
[0055] The initial endoscopic images are all endoscopic images acquired during the endoscopic examination (including endoscopic images with and without the target object), the intermediate endoscopic images are endoscopic images with the target object among all the initial endoscopic images, and the target object recognition network is used to filter out the intermediate endoscopic images from the initial endoscopic images.
[0056] Specifically, in practical applications, the target object recognition network is pre-trained multiple times using target object features (e.g., color features, shape features, etc.) to enable the trained network to recognize these features. Then, multiple initial endoscopic images are input into the trained network to extract features from each image. If these features match the target object's characteristics, the initial image is designated as an intermediate endoscopic image, and the area covered by the target object is marked within this image. Optionally, the target object recognition network is a trained YOLOv3 (YouOnly Look Once version 3) network.
[0057] For example, the first initial endoscopic image, the second initial endoscopic image, and the third initial endoscopic image are input into the trained YOLOv3 network. The YOLOv3 network identifies the initial endoscopic image features corresponding to the first initial endoscopic image as the first feature, the initial endoscopic image features corresponding to the second initial endoscopic image as the second feature, and the initial endoscopic image features corresponding to the third initial endoscopic image as the third feature. Among these, the second and third features are target object features. Therefore, the second initial endoscopic image and the third initial endoscopic image are determined to be intermediate endoscopic images.
[0058] Furthermore, the image region where the target object is located in each intermediate endoscopic image is determined. If the target object is located in a non-central region of the intermediate endoscopic image, it indicates that the intermediate endoscopic image has defects such as incomplete display of tissue parts or the target object located in the tissue parts being displayed in the edge region of the endoscopic image. Therefore, this intermediate endoscopic image is used as the endoscopic image to be processed. Optionally, the non-preset region (i.e., non-central region of the field of view) can be set according to the actual situation.
[0059] For example, the non-preset area is the pixel area in the intermediate endoscope image other than the pixels in rows 500 and columns 500 to 1000 and columns 1000, as shown in Figure 2. The target object 2001 in the first intermediate endoscope image 2003 is located in the pixel area in rows 100 and columns 200 to 300 and columns 400 of the first intermediate endoscope image 2003, while the target object in the second intermediate endoscope image is located in the pixel area in rows 600 and columns 700 to 800 and columns 900 of the second intermediate endoscope image. Therefore, the first intermediate endoscope image 2003 is determined to be the endoscope image to be processed.
[0060] S102. Input the endoscopic image to be processed into the image supplementation network to obtain the reference endoscopic image.
[0061] The reference endoscopic image is an endoscopic image that fully displays the tissue area. The image supplementation network is used to determine the reference endoscopic image corresponding to the endoscopic image to be processed. Specifically, because the tissue area displayed in the endoscopic image to be processed is incomplete and the target object is displayed in the edge area of the image, the physician cannot accurately determine the specific location of the target object in the tissue area based on the endoscopic image to be processed. Therefore, it is necessary to fully display the tissue area in the endoscopic image to adjust the position of the target object to the center of the image's field of view, thereby making it easier for the physician to determine the specific location of the target object.
[0062] In this embodiment, a GAN (Generative Adversarial Network) is used to fully display the tissue parts in the endoscopic image to be processed. Specifically, the GAN includes an image completion network. First, the endoscopic image to be processed is input into the image completion network to extract tissue part features from the endoscopic image. If a tissue part feature is a reference tissue part feature, the endoscopic image corresponding to that reference tissue part feature is determined and used as the reference endoscopic image. In practical applications, the image completion network is pre-trained extensively to acquire the ability to recognize reference tissue part features (i.e., image features of the reference endoscopic image) and has established a tissue part completion library that stores the mapping relationship between each reference tissue part feature and the reference endoscopic image (which fully displays the tissue part corresponding to the reference tissue part feature). The image completion network can determine the endoscopic image corresponding to one or more reference tissue part features by querying this tissue part completion library and use that endoscopic image as the reference endoscopic image corresponding to the endoscopic image to be processed.
[0063] For example, as shown in Figure 2, the endoscopic image to be processed (i.e., the first intermediate endoscopic image 2003) is input into an image supplementation network to extract tissue features from the endoscopic image to be processed. The tissue features are determined to be gastric antrum structural features. By querying the tissue supplementation library, the endoscopic image corresponding to the gastric antrum structural features is determined to be a complete gastric antrum structural endoscopic image 2004 (wherein, the complete gastric antrum structural endoscopic image 2004 shows a complete gastric antrum structure 2002). Therefore, the complete gastric antrum structural endoscopic image 2004 is determined to be the reference endoscopic image corresponding to the endoscopic image to be processed.
[0064] S103. Input the endoscope image to be processed and the reference endoscope image into the image generation network to generate the target endoscope image, wherein the target object is located within a preset area of the target endoscope image.
[0065] The target endoscopic image is an endoscopic image in which the target object is located within a preset area of the image (e.g., in the central region of the image's field of view) and the tissue is fully displayed. The image generation network is used to generate the target endoscopic image. Specifically, after determining the complete endoscopic image of the tissue where the target object is located (i.e., the reference endoscopic image corresponding to the target object) through the above steps, it is necessary to determine the corresponding position of the target object in the complete tissue based on the relevant information of the background mucosa in the endoscopic image to be processed. Then, the endoscopic image to be processed is combined with the reference endoscopic image to display the target object in the endoscopic image to be processed at the corresponding position in the reference endoscopic image. The corresponding background mucosa is generated at the corresponding position in the reference endoscopic image based on the relevant information of the background mucosa, and the target object is located in the central region of the field of view in the generated endoscopic image, thus enabling a more intuitive reconstruction of the specific position of the target object in the fully displayed tissue.
[0066] In this embodiment, a GAN network is used to generate a target endoscopic image where the target object is located within a preset region of the image and the tissue part is fully displayed, based on the target object and a reference endoscopic image. Specifically, the GAN network also includes an image generation network. First, the region image of the area where the target object is located is determined in the endoscopic image to be processed, and the tissue mucosal features of the region image are extracted. Then, the endoscopic image to be processed, the reference endoscopic image, and the tissue mucosal features are input into the image generation network to determine the stitching position of the target object in the reference endoscopic image. The endoscopic image to be processed and the reference endoscopic image are stitched together according to the stitching position to obtain a stitched image. At the same time, a tissue mucosal object is generated in the stitched image according to the tissue mucosal features. The stitched image with the tissue mucosal object obtained through the above steps is used as the target endoscopic image.
[0067] For example, as shown in Figure 3, a region image 2005 is determined in the endoscopic image to be processed (i.e., the first intermediate endoscopic image 2003) to identify the area where the target object 2001 is located, and the tissue mucosal features of the region image 2005 are extracted. Then, the endoscopic image to be processed, the reference endoscopic image (i.e., the endoscopic image 2004 with complete gastric antrum structure) and the tissue mucosal features are input into an image generation network to determine the splicing position of the target object 2001 in the reference endoscopic image. Based on the splicing position, the endoscopic image to be processed and the reference endoscopic image are spliced to obtain a spliced image. At the same time, a tissue mucosal object 2006 is generated in the spliced image based on the tissue mucosal features. The spliced image with the tissue mucosal object 2006 obtained through the above steps is used as the target endoscopic image 2007.
[0068] Furthermore, to avoid issues such as image ratio inconsistency, differences in background mucosal texture, and color distortion between the endoscope image to be processed and the reference endoscope image in the target endoscope image, as shown in Figure 3, in this embodiment, the endoscope image to be processed (or the reference endoscope image) in the target endoscope image 2007 is scaled to improve the connection between the two and make the image ratio more consistent, thereby making the display effect of the target endoscope image 2007 closer to the endoscope image taken in a clinical environment. Optionally, color restoration is then performed on the target endoscope image: First, the pixel values Ra, Ga, and Ba of each pixel in the endoscope image to be processed with respect to the three pixel channels are calculated, and the gain coefficient k of the three pixel channels is calculated.
[0069]
[0070] Where Raver is the average pixel value of each pixel in the endoscopic image to be processed with respect to the R channel, Gaver is the average pixel value of each pixel in the endoscopic image to be processed with respect to the G channel, and Baver is the average pixel value of each pixel in the endoscopic image to be processed with respect to the B channel.
[0071] Next, the pixel values Rb, Gb, and Bb of each pixel in the reference endoscopic image with respect to the three pixel channels are calculated, where:
[0072]
[0073]
[0074]
[0075] The pixel values of each pixel in the reference endoscope image calculated according to the above formula are used to adjust the pixel values of each pixel in the reference endoscope image to reduce the degree of tonal distortion in the target endoscope image.
[0076] Furthermore, to reduce the degree of texture difference in the target endoscopic image, in this embodiment, as shown in Figure 4, a texture region 2008 with high clarity is first selected from the endoscopic image to be processed, and the texture features of the texture region 2008 are extracted. Then, a texture object is generated based on the texture features, and the texture object is overlaid on the target endoscopic image 2007. Optionally, to further make the display effect of the target endoscopic image closer to that of endoscopic images taken in a clinical environment, super-resolution reconstruction processing is performed on the edge region of the target object in the target endoscopic image, so that the texture object has a better transition effect in the edge region of the target object, thereby improving the display effect of the target endoscopic image.
[0077] S104. Input the endoscope image to be processed and the target endoscope image into the location recognition network to obtain the target location information of the target object.
[0078] In this embodiment, target location information is used to characterize the position of the target object within the tissue, and the location recognition network is used to identify the position of the target object within the tissue. Specifically, in this embodiment, the GAN network further includes a second location recognition network, which refers to both the first and second location recognition networks. Specifically, the endoscope image to be processed is input into the first location recognition network to extract features from the endoscope image, obtaining image features of the endoscope image to be processed. Based on the image features of the endoscope image to be processed, the first location information of the target object is determined. Simultaneously, the target endoscope image is input into the second location recognition network to extract features from the target endoscope image, obtaining image features of the target endoscope image. Based on the image features of the target endoscope image, the second location information of the target object is determined. The first location information characterizes the first position of the target object relative to the tissue, and the second location information characterizes the second position of the target object relative to the tissue. Finally, the target location information of the target object is determined based on the first and second location information. Optionally, the first location recognition network is a trained deep learning network that has the ability to identify tissues.
[0079] For example, the endoscopic image to be processed is input into a first position recognition network to extract features from the endoscopic image, obtaining feature A. Based on feature A, the first position information of the target object is determined, which represents the first position of the target object relative to the gastric antrum structure. Simultaneously, the target endoscopic image is input into a second position recognition network to extract features from the target endoscopic image, obtaining feature B. Based on feature B, the second position information of the target object is determined, which represents the second position of the target object relative to the gastric antrum structure. Finally, based on the first and second position information, it is determined that the target object is located at the greater curvature of the gastric antrum (i.e., the target position information).
[0080] As described above, the image processing method provided in this application first acquires an endoscope image to be processed, wherein the target object in the endoscope image is located outside a preset region of the endoscope image. Then, the endoscope image to be processed is input into an image supplementation network to obtain a reference endoscope image. Next, the endoscope image to be processed and the reference endoscope image are input into an image generation network to generate a target endoscope image, wherein the target object is located within a preset region of the target endoscope image. Finally, the endoscope image to be processed and the target endoscope image are input into a location recognition network to obtain the target location information of the target object. Based on the endoscope image to be processed, a target endoscope image is automatically generated where the target object is located within a preset region of the field of view (e.g., the center of the field of view). The location information of the target object is accurately obtained based on its position in the endoscope image to be processed and the target endoscope image, thereby effectively improving the accuracy of target object location recognition and alleviating the current technical problem of low target object location recognition accuracy.
[0081] Based on the methods described in the above embodiments, this embodiment will be further described from the perspective of an image processing device.
[0082] Please refer to Figure 5, which specifically illustrates the image processing apparatus provided in an embodiment of this application. The image processing apparatus may include: an acquisition module 10, a reference endoscope image determination module 20, a generation module 30, and a target location information determination module 40, wherein:
[0083] (1) Obtain module 10
[0084] The acquisition module 10 is used to acquire the endoscope image to be processed; wherein the target object in the endoscope image to be processed is located in a non-preset area of the endoscope image to be processed.
[0085] Specifically, the acquisition module 10 is used for:
[0086] Multiple initial endoscopic images are input into a target object recognition network to obtain intermediate endoscopic images; wherein, the intermediate endoscopic images include initial endoscopic images containing the target object;
[0087] If the target object in the intermediate endoscope image is located in a non-preset area of the intermediate endoscope image, the intermediate endoscope image is used as the endoscope image to be processed.
[0088] Specifically, the acquisition module 10 is also used for:
[0089] Multiple initial endoscopic images are input into a target object recognition network to extract features from each initial endoscopic image, thereby obtaining the features of the initial endoscopic images.
[0090] If the features of the initial endoscopic image are the features of the target object, then the initial endoscopic image is determined to be the intermediate endoscopic image.
[0091] (2) Reference Endoscopic Image Determination Module 20
[0092] The reference endoscope image determination module 20 is used to input the endoscope image to be processed into the image supplementation network to obtain the reference endoscope image.
[0093] Specifically, the reference endoscopic image determination module 20 is used for:
[0094] The endoscopic image to be processed is input into an image supplementation network to extract tissue features from the endoscopic image and determine a reference endoscopic image based on the tissue features.
[0095] Specifically, the reference endoscope image determination module 20 is also used for:
[0096] The endoscope image to be processed is input into the image supplementation network to extract tissue features in the endoscope image to be processed. If the tissue features are reference tissue features, the endoscope image corresponding to the reference tissue features is determined.
[0097] The endoscopic image corresponding to the features of the reference tissue site is used as the reference endoscopic image.
[0098] (3) Generation module 30
[0099] The generation module 30 is used to input the endoscope image to be processed and the reference endoscope image into the image generation network to generate the target endoscope image; wherein the target object is located within a preset area of the target endoscope image.
[0100] Specifically, the generation module 30 is used for:
[0101] In the endoscopic image to be processed, determine the region image of the area where the target object is located, and extract the tissue mucosal features of the region image;
[0102] The endoscope image to be processed, the reference endoscope image, and the tissue mucosal features are input into the image generation network. The image generation network determines the stitching position of the target object in the reference endoscope image. Based on the stitching position, the endoscope image to be processed and the reference endoscope image are stitched together to obtain a stitched image. Based on the tissue mucosal features, a tissue mucosal object is generated in the stitched image.
[0103] Use a stitched image containing tissue mucosal objects as the target endoscopic image.
[0104] (4) Target location information determination module 40
[0105] The target location information determination module 40 is used to input the endoscope image to be processed and the target endoscope image into the location recognition network to obtain the target location information of the target object.
[0106] The location identification network includes a first location identification network and a second location identification network, and the target location information determination module 40 is specifically used for:
[0107] The endoscope image to be processed is input into the first location recognition network to obtain the first location information of the target object, and the target endoscope image is input into the second location recognition network to obtain the second location information of the target object;
[0108] Based on the first location information and the second location information, the target location information of the target object is determined.
[0109] In practice, the above modules can be implemented as independent entities or combined in any way to be implemented as the same or several entities. For the specific implementation of the above modules, please refer to the previous method implementation examples, which will not be repeated here.
[0110] As described above, the image processing apparatus provided in this application first acquires an endoscope image to be processed through the acquisition module 10, wherein the target object in the endoscope image to be processed is located in a non-preset area of the endoscope image to be processed. Then, the endoscope image to be processed is input into an image supplementation network through the reference endoscope image determination module 20 to obtain a reference endoscope image. After that, the endoscope image to be processed and the reference endoscope image are input into an image generation network through the generation module 30 to generate a target endoscope image, wherein the target object is located in a preset area of the target endoscope image. Finally, the target position information determination module 40 inputs the endoscope image to be processed and the target endoscope image into a position recognition network to obtain the target position information of the target object. Based on the endoscope image to be processed, a target endoscope image is automatically generated in which the target object is located in a preset area of the field of view (e.g., the center of the field of view). The target object's position information is accurately obtained based on the position of the target object in the endoscope image to be processed and the target endoscope image, thereby effectively improving the accuracy of target object position recognition and alleviating the technical problem of low target object position recognition accuracy.
[0111] Accordingly, embodiments of the present invention also provide an image processing system, including any of the image processing devices provided in the embodiments of the present invention, which can be integrated into an electronic device.
[0112] The process involves: acquiring an endoscope image to be processed, wherein the target object in the endoscope image is located outside a preset region of the endoscope image; inputting the endoscope image to be processed into an image supplementation network to obtain a reference endoscope image; inputting the endoscope image to be processed and the reference endoscope image into an image generation network to generate a target endoscope image, wherein the target object is located within a preset region of the target endoscope image; and inputting the endoscope image to be processed and the target endoscope image into a location recognition network to obtain the target location information of the target object.
[0113] The specific implementation details of each of the above devices can be found in the preceding embodiments, and will not be repeated here.
[0114] Since the image processing system can include any of the image processing devices provided in the embodiments of the present invention, it can achieve the beneficial effects that any of the image processing devices provided in the embodiments of the present invention can achieve, as detailed in the preceding embodiments, and will not be repeated here.
[0115] In addition, this application embodiment also provides an electronic device, which may be a smartphone or a computer. As shown in FIG6, the electronic device 600 includes a processor 601 and a memory 602. The processor 601 and the memory 602 are electrically connected.
[0116] The processor 601 is the control center of the electronic device 600. It connects various parts of the electronic device through various interfaces and lines. By running or loading the application program stored in the memory 602 and calling the data stored in the memory 602, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole.
[0117] In this embodiment, the processor 601 in the electronic device 600 loads the instructions corresponding to the processes of one or more applications into the memory 602 according to the following steps, and the processor 601 runs the applications stored in the memory 602 to realize various functions:
[0118] Acquire an endoscope image to be processed, wherein the target object in the endoscope image is located in a non-preset area of the endoscope image;
[0119] The endoscope image to be processed is input into the image supplementation network to obtain the reference endoscope image;
[0120] The endoscope image to be processed and the reference endoscope image are input into the image generation network to generate the target endoscope image, wherein the target object is located within a preset region of the target endoscope image;
[0121] The endoscope image to be processed and the target endoscope image are input into the location recognition network to obtain the target location information of the target object.
[0122] Figure 7 shows a specific structural block diagram of the electronic device provided in the embodiment of the present invention. The electronic device can be used to implement the image processing method provided in the above embodiment.
[0123] RF circuit 710 is used to receive and transmit electromagnetic waves, converting electromagnetic waves into electrical signals and vice versa, thereby enabling communication with communication networks or other devices. RF circuit 710 may include various existing circuit elements used to perform these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, Subscriber Identity Module (SIM) cards, memory, etc. RF circuit 710 can communicate with various networks such as the Internet, corporate intranets, and wireless networks, or communicate with other devices via wireless networks. The aforementioned wireless networks may include cellular telephone networks, wireless local area networks (WLANs), or metropolitan area networks (MANs). The aforementioned wireless networks may use various communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communication (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (WCDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Wireless Fidelity (Wi-Fi) (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), Voice over Internet Protocol (VoIP), Worldwide Interoperability for Microwave Access (Wi-Max), other protocols for email, instant messaging, and short messages, and any other suitable communication protocols, including those that have not yet been developed.
[0124] The memory 720 can be used to store software programs and modules. The processor 780 executes various functional applications and data processing by running the software programs and modules stored in the memory 720, thereby realizing the function of storing 5G capability information. The memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 720 may further include memory remotely located relative to the processor 780, and these remote memories can be connected to the electronic device 700 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0125] The input unit 730 can be used to receive input digital or character information, and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. Specifically, the input unit 730 may include a touch-sensitive surface 731 and other input devices 732. The touch-sensitive surface 731, also known as a touch display screen or touchpad, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch-sensitive surface 731), and drive the corresponding connection device according to a pre-set program. Optionally, the touch-sensitive surface 731 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to the processor 780, and can receive and execute commands sent by the processor 780. In addition, the touch-sensitive surface 731 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch-sensitive surface 731, the input unit 730 may also include other input devices 732. Specifically, other input devices 732 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0126] Display unit 740 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of electronic device 700. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Display unit 740 may include display panel 741, which may optionally be configured as an LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), or similar display panel 741. Further, touch-sensitive surface 731 may cover display panel 741. When touch-sensitive surface 731 detects a touch operation on or near it, it transmits the information to processor 780 to determine the type of touch event. Subsequently, processor 780 provides corresponding visual output on display panel 741 according to the type of touch event. Although in FIG. 7, touch-sensitive surface 731 and display panel 741 are implemented as two separate components to realize input and output functions, in some embodiments, touch-sensitive surface 731 and display panel 741 can be integrated to realize input and output functions.
[0127] The electronic device 700 may also include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 741 according to the ambient light level, and the proximity sensor can turn off the display panel 741 and / or backlight when the electronic device 700 is moved to the ear. As a type of motion sensor, a gravity acceleration sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometers, taps), etc. Other sensors that the electronic device 700 may also be equipped with, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0128] Audio circuitry 760, speaker 761, and microphone 762 provide an audio interface between the user and electronic device 700. Audio circuitry 760 converts received audio data into electrical signals and transmits them to speaker 761, where speaker 761 converts them into sound signals for output. Conversely, microphone 762 converts collected sound signals into electrical signals, which are then received by audio circuitry 760, converted back into audio data, and processed by processor 780. The audio data is then transmitted via RF circuitry 710 to, for example, another terminal, or output to memory 720 for further processing. Audio circuitry 760 may also include an earphone jack to facilitate communication between peripheral headphones and electronic device 700.
[0129] Electronic device 700, through transmission module 770 (e.g., Wi-Fi module), enables users to send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 7 shows transmission module 770, it is understood that it is not an essential component of electronic device 700 and can be omitted as needed without altering the essence of the invention.
[0130] The processor 780 is the control center of the electronic device 700. It connects to various parts of the phone via various interfaces and lines, and performs various functions and processes data of the electronic device 700 by running or executing software programs and / or modules stored in the memory 720, and by calling data stored in the memory 720. Optionally, the processor 780 may include one or more processing cores; in some embodiments, the processor 780 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 780.
[0131] The electronic device 700 also includes a power supply 790 (such as a battery) that supplies power to various components. In some embodiments, the power supply may be logically connected to the processor 780 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. The power supply 790 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0132] Although not shown, the electronic device 700 may also include a camera (such as a front-facing camera and a rear-facing camera), a Bluetooth module, etc., which will not be described in detail here. Specifically, in this embodiment, the display unit of the electronic device is a touch screen display, and the electronic device also includes a memory and one or more programs, wherein one or more programs are stored in the memory and configured to be executed by one or more processors. One or more programs contain instructions for performing the following operations:
[0133] Acquire an endoscope image to be processed, wherein the target object in the endoscope image is located in a non-preset area of the endoscope image;
[0134] The endoscope image to be processed is input into the image supplementation network to obtain the reference endoscope image;
[0135] The endoscope image to be processed and the reference endoscope image are input into the image generation network to generate the target endoscope image, wherein the target object is located within a preset region of the target endoscope image;
[0136] The endoscope image to be processed and the target endoscope image are input into the location recognition network to obtain the target location information of the target object.
[0137] In practice, the above modules can be implemented as independent entities or combined in any way to be implemented as the same or several entities. For the specific implementation of the above modules, please refer to the previous method implementation examples, which will not be repeated here.
[0138] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. Therefore, embodiments of the present invention provide a storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps in any of the image processing methods provided in the embodiments of the present invention.
[0139] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0140] Since the instructions stored in the storage medium can execute the steps of any of the image processing methods provided in the embodiments of the present invention, the beneficial effects that any of the image processing methods provided in the embodiments of the present invention can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0141] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0142] In summary, although the present application has disclosed the preferred embodiments as described above, the above preferred embodiments are not intended to limit the present application. Those skilled in the art can make various modifications and refinements without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be determined by the scope defined in the claims.
Claims
1. An image processing method, characterized in that, include: Acquire an endoscopic image to be processed; wherein the target object in the endoscopic image to be processed is located within a non-preset region of the endoscopic image to be processed; input the endoscopic image to be processed into an image supplementation network to obtain a reference endoscopic image; input the endoscopic image to be processed and the reference endoscopic image into an image generation network to generate a target endoscopic image; wherein the target object is located within a preset region of the target endoscopic image; input the endoscopic image to be processed and the target endoscopic image into a location recognition network to obtain target location information of the target object; wherein the step of inputting the endoscopic image to be processed into an image supplementation network to obtain a reference endoscopic image includes: inputting the endoscopic image to be processed into the image supplementation network to extract tissue features in the endoscopic image to be processed through the image supplementation network, and determining the reference endoscopic image based on the tissue features; the step of inputting the endoscopic image to be processed into the image supplementation network to extract tissue features in the endoscopic image to be processed through the image supplementation network, and determining the reference endoscopic image based on the tissue features includes: inputting the endoscopic image to be processed into the image supplementation network to extract tissue features in the endoscopic image to be processed through the image supplementation network, and determining the reference endoscopic image based on the tissue features, includes: inputting the endoscopic image to be processed into the image supplementation network to extract tissue features in the endoscopic image to be processed through the image supplementation network, and determining the reference endoscopic image based on the tissue features; The process involves inputting the endoscopic image to the image supplementation network to extract tissue features from the endoscopic image to be processed. If the tissue feature is a reference tissue feature, an endoscopic image corresponding to the reference tissue feature is determined. The endoscopic image corresponding to the reference tissue feature is then used as the reference endoscopic image. The step of inputting the endoscopic image to be processed and the reference endoscopic image into an image generation network to generate a target endoscopic image includes: determining a region image of the area where the target object is located in the endoscopic image to be processed, and extracting tissue mucosal features from the region image; inputting the endoscopic image to be processed, the reference endoscopic image, and the tissue mucosal features into the image generation network to determine the stitching position of the target object in the reference endoscopic image, and stitching the endoscopic image to be processed and the reference endoscopic image according to the stitching position to obtain a stitched image, and generating a tissue mucosal object in the stitched image according to the tissue mucosal features; and using the stitched image containing the tissue mucosal object as the target endoscopic image.
2. The image processing method according to claim 1, characterized in that, The step of acquiring the endoscope image to be processed includes: inputting multiple initial endoscope images into a target object recognition network to obtain an intermediate endoscope image; wherein, the intermediate endoscope image includes the initial endoscope image containing the target object; if the target object in the intermediate endoscope image is located in the non-preset region of the intermediate endoscope image, the intermediate endoscope image is used as the endoscope image to be processed.
3. The image processing method according to claim 2, characterized in that, The step of inputting multiple initial endoscopic images into a target object recognition network to obtain an intermediate endoscopic image includes: inputting multiple initial endoscopic images into the target object recognition network to extract features from each initial endoscopic image through the target object recognition network to obtain initial endoscopic image features; if the initial endoscopic image features are target object features, determining the initial endoscopic image as the intermediate endoscopic image.
4. The image processing method according to claim 1, characterized in that, The location recognition network includes a first location recognition network and a second location recognition network. The step of inputting the endoscope image to be processed and the target endoscope image into the location recognition network to obtain the target location information of the target object includes: inputting the endoscope image to be processed into the first location recognition network to obtain the first location information of the target object, and inputting the target endoscope image into the second location recognition network to obtain the second location information of the target object; and determining the target location information of the target object based on the first location information and the second location information.
5. An image processing apparatus, characterized in that, include: An acquisition module is used to acquire an endoscopic image to be processed; wherein the target object in the endoscopic image to be processed is located within a non-preset region of the endoscopic image to be processed; a reference endoscopic image determination module is used to input the endoscopic image to be processed into an image supplementation network to obtain a reference endoscopic image; a generation module is used to input the endoscopic image to be processed and the reference endoscopic image into an image generation network to generate a target endoscopic image; wherein the target object is located within a preset region of the target endoscopic image; a target location information determination module is used to input the endoscopic image to be processed and the target endoscopic image into a location recognition network to obtain the target location information of the target object; wherein the reference endoscopic image determination module is further used to input the endoscopic image to be processed into the image supplementation network to extract tissue features in the endoscopic image to be processed through the image supplementation network, and to determine the reference endoscopic image based on the tissue features; the reference endoscopic image determination module is further used to input the endoscopic image to be processed into the image supplementation network to extract tissue features in the endoscopic image to be processed, and to determine the reference endoscopic image based on the tissue features; the reference endoscopic image determination module is further used to input the endoscopic image to be processed into the image supplementation network to obtain a target location information of the target object; The image supplementation network is input to extract tissue features from the endoscopic image to be processed. If the tissue feature is a reference tissue feature, an endoscopic image corresponding to the reference tissue feature is determined. The endoscopic image corresponding to the reference tissue feature is used as the reference endoscopic image. The generation module is further used to determine the region image of the area where the target object is located in the endoscopic image to be processed, and extract the tissue mucosal features of the region image. The endoscopic image to be processed, the reference endoscopic image, and the tissue mucosal features are input to the image generation network to determine the stitching position of the target object in the reference endoscopic image. The endoscopic image to be processed and the reference endoscopic image are stitched together according to the stitching position to obtain a stitched image. A tissue mucosal object is generated in the stitched image according to the tissue mucosal features. The stitched image containing the tissue mucosal object is used as the target endoscopic image.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the image processing method according to any one of claims 1 to 4.
7. An electronic device, characterized in that, The method includes a processor and a memory, the processor being electrically connected to the memory, the memory being used to store instructions and data, and the processor being used to execute the steps of the image processing method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Endoscope image processing method and device, electronic equipment and storage medium
CN112712515A
Endoscopic image processing method and device, endoscope and storage medium
CN113763298A