AI-based nose and throat lesion image recognition method and system

By using an AI-based image recognition method for nasopharyngeal lesions, combining grayscale and color images, and performing image cross-correlation, registration, and modal fusion, the problem of low accuracy in identifying nasopharyngeal disease types in existing technologies is solved, achieving full coverage detection and accurate diagnosis of lesions in the nasopharynx and larynx.

CN121304643APending Publication Date: 2026-01-09AFFILIATED ZHONGSHAN HOSPITAL OF DALIAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511647139.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Current imaging technology cannot diagnose all types of nasopharyngeal diseases in patients based on imaging results, resulting in low accuracy in identifying nasopharyngeal disease types.

Method used

An AI-based image recognition method for nasopharyngeal lesions is adopted. By acquiring grayscale and color images, image cross-correlation and registration are performed to extract similar and different features. Modal fusion is then carried out, and a fused image is generated using a neural network model to achieve accurate identification of lesions in the nasopharynx.

Benefits of technology

It improves the accuracy of identifying nasopharyngeal disease types, enables comprehensive detection of lesion types, enhances the ability to identify both surface and internal lesions, and improves diagnostic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121304643A_ABST
    Figure CN121304643A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical monitoring, in particular to an AI-based nose and throat disease change image recognition method and system, and the method comprises the steps: carrying out the image switching of an area where a nose and throat part in a first gray image is located and an area where the nose and throat part in a first color image is located, and generating a second gray image and a second color image; generating a third grayscale image and a third color image through registration processing; extracting same features and difference features of the nose and throat connection part from the third gray level image and the third color image, and performing modal fusion processing to generate a fusion image; and judging the lesion type and lesion cause of the nose and throat part by combining the extracted same features and difference features, thereby realizing the detection of the lesion of the nose and throat part. The nose and throat disease change image recognition system can execute the nose and throat disease change image recognition method. Through the arrangement, the accuracy of nose and throat disease type detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of medical technology, and more particularly to an AI-based method and system for recognizing nasopharyngeal lesions in images. Background Technology

[0002] Nasopharyngeal diseases (NLDs) are a collective term for all diseases occurring in the area connecting the nasal cavity, pharynx, and larynx. Because this area is a crucial passageway for breathing, swallowing, and vocalization, there are numerous types of NLDs. In related technologies, diagnosing NLDs generally requires imaging techniques to obtain images of the area connecting the nasal cavity, pharynx, and larynx. The specific type of NLD is determined based on the imaging results. Current imaging techniques used for detecting NLDs include conventional color endoscopy, narrow-band imaging, CT scans, magnetic resonance imaging (MRI), and computed tomography (CT). These different types of imaging techniques can detect the type of NLD in different patients. However, due to the limitations of these different imaging techniques in detecting different types of NLDs, it is impossible to diagnose all types of NLDs in a patient based solely on imaging results, thus reducing the accuracy of NLD type identification. Summary of the Invention

[0003] In view of the technical problems mentioned in the background section, an AI-based image recognition method and system for nasopharyngeal lesions is provided.

[0004] The technical means employed in this invention are as follows: An AI-based image recognition method for nasopharyngeal lesions includes the following steps: A first grayscale image and a first color image containing the nasopharynx are acquired, and the regions containing the nasopharynx in the first grayscale image and the regions containing the nasopharynx in the first color image are cross-correlated to generate a second grayscale image and a second color image; the regions containing the nasopharynx in the second grayscale image and the regions containing the nasopharynx in the second color image are roughly aligned in space. For the nasopharyngeal junction in the second grayscale image, a registration process is performed with the nasopharyngeal junction in the second color image. Through the registration process, it is ensured that the nasopharyngeal junctions with the same spatial position in the second grayscale image and the second color image form a corresponding relationship, and a third grayscale image and a third color image are generated; the spatial position of the same nasopharyngeal junction in the third grayscale image and the third color image is completely consistent. Extract the common and different features of the nasopharyngeal junction from the third grayscale image and the third color image; Modality fusion processing is performed on the extracted identical features and differential features to merge the identical features and differential features, thereby generating a fusion map; Based on the fusion map and the similar and dissimilar features, the lesion sites in the nasopharynx are identified.

[0005] Furthermore, the device that acquires the first grayscale image is defined as a grayscale device, and the device that acquires the first color image is defined as a color image device. The image cross-correlation includes: The height, pixel pitch, and focal length of the grayscale device and the color image device are obtained, and the width and height data of the first color image and the first grayscale image are obtained simultaneously. Based on the height, pixel pitch, and focal length of the acquired grayscale and color image devices, the resolution quotient of the grayscale and color image devices is calculated. Combining the calculated resolution quotient with the width and height information of the acquired first color image and first grayscale image, the first resampling factor corresponding to the first color image is determined; A second resampling factor is determined based on the first resampling factor, and the first color image is resampled using the second resampling factor to generate a resampled color image. Image cross-correlation processing is performed on the region where the nasopharynx is located in the first grayscale image and the region where the nasopharynx is located in the resampled color image.

[0006] Furthermore, the nasopharyngeal connection region in the second grayscale image is defined as the first nasopharyngeal connection region, and the nasopharyngeal connection region in the second color image is defined as the second nasopharyngeal connection region. The registration of the nasopharyngeal junction in the second grayscale image and the nasopharyngeal junction in the second color image includes the following steps: Obtain the spatial location information of the first nasopharyngeal junction and the second nasopharyngeal junction; Based on the obtained spatial location, the intersection area and the union area between the first nasopharyngeal connection point and its corresponding second nasopharyngeal connection point are calculated respectively. Then, the area intersection-union ratio between the first nasopharyngeal connection point and the corresponding second nasopharyngeal connection point is calculated using the intersection area and the union area. Filter and remove the first nasopharyngeal junction and its corresponding second nasopharyngeal junction with an area intersection ratio less than a preset threshold.

[0007] Furthermore, obtaining the same features of the nasopharyngeal junction in the third grayscale image and the third color image includes the following steps: Obtain the first image feature corresponding to the nasopharyngeal connection in the third grayscale image, and the second image feature corresponding to the nasopharyngeal connection in the third color image; Perform a cross-product operation on the obtained first image features and second image features.

[0008] Furthermore, obtaining the difference features of the nasopharyngeal junction in the third grayscale image and the third color image includes the following steps: Obtain the first image features of the nasopharynx-laryngeal junction in the third grayscale image, and the second image features of the nasopharynx-laryngeal junction in the third color image; The first image feature is subtracted from the second image feature, and the first difference feature map is generated by this operation; Image processing operations are performed on the generated first difference feature map, and attention features are determined based on the processing results; The product of the acquired attention features and the first image features is calculated, and a second difference feature map is generated by the product. The second difference feature map contains the difference features in the third grayscale image features, thereby realizing the acquisition of the difference features.

[0009] Furthermore, the generation of the fused graph includes the following steps: The same and different features are concatenated. Based on a single convolution operation, the concatenated same and different features are aligned to generate a fused image.

[0010] Furthermore, the model for generating the fused graph is defined as a neural network model, and the generation of the fused graph includes: The third grayscale image and the third color image are input into the neural network model; Under the action of the neural network model, the third grayscale image and the third color image are sequentially subjected to operations such as convolution normalization, feature extraction, adaptive convolution, structural reparameterization, fusion of data residual structures, and feature concatenation, thereby generating the same features and different features through these operations. The obtained identical feature maps and differing features are spliced ​​together to generate a fused map.

[0011] Furthermore, the image cross-correlation of the nasopharyngeal region in the first grayscale image and the nasopharyngeal region in the resampled color image also includes the following steps: A first boundary circle is set along the contour of the nasopharynx in the first grayscale image; at the same time, a second boundary circle is set along the contour of the nasopharynx in the resampled color image. The first center coordinates of the nasopharynx in the first grayscale image, the second center coordinates of the nasopharynx in the resampled color image, the width and height of the first boundary circle, and the width and height of the second boundary circle are obtained, and the relative positions of the nasopharynx in the first grayscale image and the resampled color image are matched accordingly. Based on the relative positions of the nasopharynx and larynx in the first grayscale image and the resampled color image, image cross-correlation processing is performed on the regions where the nasopharynx and larynx are located in the first grayscale image and the regions where the nasopharynx and larynx are located in the resampled color image.

[0012] Further, the nasopharyngeal junction in the second grayscale image is defined as the first nasopharyngeal junction, and the nasopharyngeal junction in the second color image is defined as the second nasopharyngeal junction. A third boundary circle is set along the contour of the first nasopharyngeal junction, and a fourth boundary circle is set along the contour of the second nasopharyngeal junction. The registration of the nasopharyngeal junction in the second grayscale image and the color image also includes the following steps: Based on the area intersection-union ratio of the first nasopharyngeal connection and the second nasopharyngeal connection, the first center coordinate, the second center coordinate, the width and height of the third boundary circle, and the width and height of the fourth boundary circle, a registration operation is performed on the first nasopharyngeal connection and the second nasopharyngeal connection, which are spatially closest in the second grayscale image and the second color image, and a nasopharyngeal connection registration component is generated through this registration. The registration components of the nasopharyngeal connection area with an area intersection ratio less than a preset threshold are filtered out and removed.

[0013] The present invention also provides an AI-based image recognition system for nasopharyngeal diseases, including an acquisition device and an image recognition device for nasopharyngeal diseases. The acquisition device is used to acquire a first grayscale image and a first color image; the image recognition device for nasopharyngeal diseases is used to execute an image recognition method for nasopharyngeal diseases.

[0014] Compared with the prior art, the present invention has the following advantages: 1. Current diagnostic techniques for nasopharyngeal diseases, such as conventional color endoscopy and CT scans, have significant limitations. For example, color endoscopy can only display surface features and is difficult to detect internal lesions, while CT scans, although capable of displaying internal structures, are insufficient for identifying subtle surface lesions, leading to an inability to diagnose all types of lesions and reduced accuracy. This application, however, combines a first grayscale image with a first color image to achieve comprehensive lesion type detection through multimodal collaboration.

[0015] 2. Image cross-correlation can improve the alignment accuracy of grayscale and color images. Furthermore, registration processing incorporates the area intersection-union ratio (IU / R) and cost function to optimize the registration effect. The intersection and union areas of the first and second nasopharyngeal junctions are calculated to obtain the IU / R. Registration components with IU / R ratios less than a preset threshold are eliminated, ensuring accurate correspondence between spatially identical regions in the grayscale and color images. Simultaneously, the registration process is transformed into a task assignment problem of solving the minimum cost function, selecting the registration combination with the largest IU / R to generate a third grayscale image and a third color image with completely identical spatial positions, thereby improving the accuracy of the image data for feature extraction.

[0016] 3. In feature extraction and modality fusion, effective features are enhanced through techniques such as cross-product and attention mechanisms. Cross-product operations are performed on the first and second image features to enhance identical features and suppress dissimilar features, improving the accuracy of surface lesion recognition. A first dissimilar feature map is generated by image subtraction, and then processed through average pooling, max pooling, and a shared multilayer perceptron to obtain attention features. These attention features are then multiplied with the first dissimilar feature map to generate a second dissimilar feature map, highlighting internal lesion features. Finally, convolution is used to align identical and dissimilar features, generating a fusion map. The fusion map, combined with the two types of features, determines the lesion type and cause. These multi-stage technical optimizations collectively ensure the accuracy of lesion detection. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a structural block diagram of the nasopharyngeal lesion image recognition device according to an embodiment of this application.

[0019] Figure 2 This is a flowchart of the nasopharyngeal lesion image recognition method according to an embodiment of this application.

[0020] Figure 3 This is a flowchart illustrating step S1 in an embodiment of this application.

[0021] Figure 4 This is a flowchart illustrating step S15 in an embodiment of this application.

[0022] Figure 5 This is a flowchart illustrating the first step S2 in an embodiment of this application.

[0023] Figure 6This is a flowchart illustrating the second step S21 in an embodiment of this application.

[0024] Figure 7 This is a flowchart illustrating the first step S3 in an embodiment of this application.

[0025] Figure 8 This is a flowchart illustrating the second step S3 in an embodiment of this application.

[0026] Figure 9 This is a flowchart illustrating step S4 in an embodiment of this application.

[0027] Figure 10 This is a structural block diagram of the nasopharyngeal lesion image recognition system according to an embodiment of this application. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] The method embodiments provided in this example can be executed in the nasopharyngeal disease image recognition device 100, the nasopharyngeal disease image recognition system 200, or a similar device. Figure 1 This is a hardware structure block diagram of a nasopharyngeal disease image recognition device 100 implementing an embodiment of this application. For example... Figure 1 As shown, the nasopharyngeal disease image recognition device 100 may include one or more ( Figure 1 (Only one is shown in the image) Memory 12 and processor 11.

[0031] The memory 12 stores program instructions, such as application software programs and modules, like the computer program corresponding to an AI-based nasopharyngeal lesion image recognition method in this embodiment. The processor 11 executes various functional applications and data processing by running the computer program stored in the memory 12, thus implementing the aforementioned method. The memory 12 may include high-speed random access memory (RAM) 12, and may also include non-volatile memory 12, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory 12. In some instances, the memory 12 may further include remotely located memory 12 relative to the processor 11, which can be connected to the nasopharyngeal lesion image recognition device 100 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0032] The processor 11 is used to execute program instructions stored in the memory 12. The processor 11 may include, but is not limited to, a microprocessor (MCU) or a programmable gate array (FPGA). The aforementioned nasopharyngeal disease image recognition device 100 may also include a transmission device 13 for communication functions and an input / output device 14. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the nasopharyngeal disease image recognition device 100 described above. For example, the nasopharyngeal disease image recognition device 100 may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown are illustrated.

[0033] The transmission device 13 is used to receive or send data via a network. This network includes a wireless network provided by the communication provider of the nasopharyngeal disease image recognition device 100. In one example, the transmission device 13 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 13 can be a Radio Frequency (RF) module for wireless communication with the Internet.

[0034] This embodiment provides an AI-based image recognition method for nasopharyngeal diseases. Figure 2This is a flowchart of an AI-based image recognition method for nasopharyngeal diseases according to an embodiment of this application. This lesion detection method is used to perform fusion modality lesion detection in the nasopharyngeal region.

[0035] In this application, the nasopharyngeal region includes the area where the nasopharynx and larynx meet. like Figure 2 As shown, the image recognition method for nasopharyngeal diseases provided in this application includes the following steps: Step S1: Obtain a first grayscale image and a first color image containing the nasopharynx, and perform image cross-correlation on the nasopharynx region in the first grayscale image and the nasopharynx region in the first color image to generate a second grayscale image and a second color image. The nasopharynx region in the second grayscale image and the nasopharynx region in the second color image are roughly aligned in space.

[0036] In this application, by acquiring a first grayscale image containing the nasopharynx, the internal state of the nasopharynx can be analyzed; simultaneously, by acquiring a first color image containing the nasopharynx, the surface state of the nasopharynx can be analyzed. This setup facilitates the subsequent detection of different types of lesions in the nasopharynx, thereby improving the accuracy of lesion detection by the intelligent nasopharyngeal recognition method.

[0037] Image cross-correlation is a mathematical operation that measures the similarity between two images (or two regions of images) at a specific relative position. Since the relative positions of the nasopharynx and larynx differ between the first grayscale image and the first color image, in the above setup, image cross-correlation can be used to adjust their relative positions, ensuring consistency between the nasopharynx and larynx region in the second grayscale image and the second color image. This adjustment facilitates comparison of identical areas of the nasopharynx and larynx in the second grayscale and second color images, thereby supporting improved registration efficiency for the connected nasopharynx and larynx.

[0038] Specifically, such as Figure 3 As shown, in this application, the device for acquiring the first grayscale image is defined as a grayscale device, and the device for acquiring the first color image is defined as a color image device. The image cross-correlation step S1 includes the following steps: Step S11: Obtain the height, pixel pitch, and device focal length of the grayscale device and the color image device, and obtain the width and height of the first color image and the first grayscale image. In this application, the pixel pitch is the physical distance between the center points of two adjacent pixels on the grayscale device or the color image device.

[0039] Step S12: Calculate the resolution quotient of grayscale devices and color image devices based on height, pixel pitch, and device focal length.

[0040] Because grayscale devices and color imaging devices differ in their optical characteristics, this difference leads to a difference in their spatial resolution. Calculating the resolution quotient between grayscale and color imaging devices provides necessary support for calculating the first resampling factor of the first color image, thus facilitating subsequent image preprocessing. The formula for calculating the resolution quotient is as follows: ; in, For the spatial resolution of color imaging devices, The height of the color imaging device when the first color image is captured. For the pixel pitch of color imaging devices, The focal length of a color imaging device, This refers to the spatial resolution of grayscale devices. The height of the grayscale device when capturing the first grayscale image. This refers to the pixel pitch of a grayscale device. The focal length of the grayscale device. This refers to the resolution quotient of grayscale devices and color image devices. In other words, the resolution quotient of grayscale devices and color image devices can be calculated using the formula for calculating the resolution quotient.

[0041] In this embodiment, the nasopharynx can be imaged using a grayscale imaging device (e.g., a magnetic resonance imaging device) or a color imaging device (endoscope), thereby acquiring first grayscale images and first color images of the nasopharynx at different locations, which helps to improve the efficiency of acquiring the first grayscale images and first color images. This application does not limit this.

[0042] Step S13: Based on the resolution quotient, the width and height of the first color image and the first grayscale image, determine the first resampling factor of the first color image. The formula for calculating the first resampling factor in the above settings is as follows: ; ; in, The first resampling factor is the first resampling factor in the width direction of the first color image. The first resampling factor in the height direction of the first color image. The width of the first color image, The width of the first grayscale image. The height of the first color image, Let be the height of the first grayscale image. That is, the first resampling factor of the first color image can be calculated in both the width and height directions using the formula for calculating the first resampling factor.

[0043] Step S14: Based on the first resampling factor, obtain the second resampling factor, and resample the first color image to generate a resampled color image.

[0044] In the above setup, based on the first resampling factor, multiple search operations can be performed on the first color image using a local grid. During any search, a search is conducted in the vicinity of the first resampling factor using a local grid, and the second resampling factor corresponding to that search is obtained. Then, the first color image is resampled based on this second resampling factor to generate a resampled color image. After resampling, the resampled color image and the first grayscale image are cross-correlated, and the correlation coefficient corresponding to this cross-correlation operation is calculated. When all search operations are completed, the correlation coefficient values ​​obtained from different searches are compared, and the second resampling factor corresponding to the largest matching correlation coefficient is the optimal second resampling factor. The formula for calculating the second resampling factor during any search is as follows: ; in, This is the second resampling factor in the width direction of the resampled color image. To set a constant, This is the second resampling factor in the height direction of the resampled color image. The second resampling factor can be calculated using the formula for calculating the second resampling factor.

[0045] Based on the aforementioned optimal second resampling factor, the most similar first grayscale image and resampled color image of the nasopharyngeal region can be selected, thereby improving the accuracy of image cross-correlation.

[0046] Step S15: Perform image cross-correlation between the nasopharyngeal region in the first grayscale image and the nasopharyngeal region in the resampled color image.

[0047] In the above setup, the first color image is first resampled based on the optimal second resampling factor to obtain a resampled color image. Next, an image cross-correlation operation is performed between the resampled color image and the first grayscale image, yielding multiple cross-correlation results. Then, from these cross-correlation results, the set of results showing the closest similarity between the nasopharyngeal region in the first grayscale image and the resampled color image is selected and used for subsequent registration of the nasopharyngeal region. This operation eliminates the viewing angle difference between the first grayscale image and the first color image, thus helping to avoid the adverse effects of the viewing angle difference on the image cross-correlation operation between the first grayscale image and the first color image, ultimately improving the accuracy of the image cross-correlation.

[0048] More specifically, such as Figure 4 As shown, step S15, which involves performing image cross-correlation between the nasopharyngeal region in the first grayscale image and the nasopharyngeal region in the resampled color image, further includes: Step S151: Set a first boundary circle along the contour of the nasopharynx in the first grayscale image, and set a second boundary circle along the contour of the nasopharynx in the resampled color image. This setting improves the clarity of the nasopharynx in both the first grayscale image and the resampled color image. This improved clarity facilitates subsequent image cross-correlation operations on the regions containing the nasopharynx in both the first grayscale image and the resampled color image.

[0049] Step S152: Obtain the first center coordinates of the nasopharynx in the first grayscale image, the second center coordinates of the nasopharynx in the resampled color image, the width and height of the first boundary circle, and the width and height of the second boundary circle, and match the relative position of the nasopharynx in the first grayscale image and the resampled color image. In the above settings, coordinates are used to determine the relative position of the nasopharynx in the first grayscale image and the resampled color image. This method improves the accuracy of determining the relative position of the nasopharynx in these two images. The improved accuracy of the relative position determination can further improve the accuracy of subsequent image cross-correlation operations on the first grayscale image and the resampled color image.

[0050] Step S153: Perform image cross-correlation on the regions containing the nasopharynx and larynx in the first grayscale image and the resampled color image based on their relative positions in the first grayscale image and the resampled color image. This setting ensures that the spatial position of the nasopharynx and larynx in the second grayscale image and the second color image is more consistent. This improved positional consistency facilitates subsequent registration of the connected nasopharyngeal regions, thereby improving the efficiency of the registration process.

[0051] like Figure 2 As shown, step S2: For the nasopharyngeal connection part in the second grayscale image, a registration process is performed with the nasopharyngeal connection part in the second color image. Through the above registration process, it is ensured that the nasopharyngeal connection parts with the same spatial position in the second grayscale image and the second color image form a corresponding relationship. Based on this, a third grayscale image and a third color image are generated. According to the obtained third grayscale image and third color image, the spatial position of the same nasopharyngeal connection parts is completely consistent.

[0052] The above settings facilitate the acquisition of image features of the nasopharyngeal junction, which is spatially located in both the second grayscale and second color images. Furthermore, comparing the image features of this nasopharyngeal junction in the second grayscale and second color images creates favorable conditions for subsequently acquiring all types of lesions present in the nasopharyngeal junction, ultimately improving the accuracy of lesion detection methods.

[0053] like Figure 5 As shown, the nasopharyngeal junction in the second grayscale image is defined as the first nasopharyngeal junction, and the nasopharyngeal junction in the second color image is defined as the second nasopharyngeal junction. Step S2: Registration of the nasopharyngeal junction in the second grayscale image and the nasopharyngeal junction in the second color image includes: Step S21: Obtain the spatial location of the first nasopharyngeal junction and the second nasopharyngeal junction.

[0054] Step S22: Based on spatial location, calculate the intersection area and the union area of ​​the first nasopharyngeal connection site and the corresponding second nasopharyngeal connection site, and calculate the area intersection-union ratio of the first nasopharyngeal connection site and the corresponding second nasopharyngeal connection site.

[0055] In this application, the Intersection over Union (IoU) ratio is a metric used in computer vision to measure the degree of overlap between two regions. It is typically calculated by taking the ratio of the intersection area to the union area of ​​the predicted and real regions, thus quantifying the matching degree between the two regions. Since the second grayscale image and the second color image have undergone image cross-correlation, this operation allows the second infrared image and the second grayscale image to overlap on the same plane. Based on this, the IoU ratio between the first and second nasopharyngeal junctions can be calculated to determine whether the calculated first and second nasopharyngeal junctions belong to the nasopharyngeal junctions with the same spatial location. Specifically, the larger the IoU ratio between the first and second nasopharyngeal junctions, the higher the degree of overlap between the two regions, and thus the greater the probability that the images in the second grayscale image and the second color image are nasopharyngeal junctions with the same spatial location in the nasopharyngeal region. The above judgment can facilitate the subsequent registration operation between the first nasopharyngeal junction and the corresponding second nasopharyngeal junction.

[0056] Step S23: Remove the first nasopharyngeal connection part and the corresponding second nasopharyngeal connection part whose area intersection ratio is less than a preset threshold, so that the first nasopharyngeal connection parts and the second nasopharyngeal connection parts with the same spatial position in the second grayscale image and the second color image correspond to each other.

[0057] In the above settings, if the area overlap ratio between the first nasopharyngeal connection and the corresponding second nasopharyngeal connection is lower than a preset threshold, it indicates that the overlap between the first and second nasopharyngeal connection is low. Therefore, it can be determined that, within the overall range of the nasopharynx, the first and second nasopharyngeal connection are not nasopharyngeal connection points located in the same spatial position. Based on this determination, the first and corresponding second nasopharyngeal connection points can be filtered out, thereby creating favorable conditions for improving the efficiency of nasopharyngeal connection point registration operations.

[0058] In some embodiments, the preset threshold can be set to 75%. Specifically, if the area overlap ratio between the first nasopharyngeal connection and the second nasopharyngeal connection is less than 75%, it can be inferred that the spatial position of the first nasopharyngeal connection in the second grayscale image does not coincide with the spatial position of the second nasopharyngeal connection in the second color image; if the area overlap ratio between the first nasopharyngeal connection and the second nasopharyngeal connection is greater than 75%, it can be inferred that the spatial position of the first nasopharyngeal connection in the second grayscale image coincides with the spatial position of the second nasopharyngeal connection in the second color image.

[0059] For example, a third boundary ring is provided along the contour of the first nasopharyngeal junction, and a fourth boundary ring is provided along the contour of the second nasopharyngeal junction, such as... Figure 6 As shown, step S21: obtaining the spatial positions of the first nasopharyngeal junction and the second nasopharyngeal junction further includes: S21a: Based on the area intersection-union ratio of the first nasopharyngeal junction and the second nasopharyngeal junction, the first center coordinate, the second center coordinate, the width and height of the third boundary circle, and the width and height of the fourth boundary circle, register the first nasopharyngeal junction and the second nasopharyngeal junction that are spatially closest in the second grayscale image and the second color image to generate a nasopharyngeal junction registration component.

[0060] In this application, the first boundary circle includes the boundaries of all first nasopharyngeal junctions in the second grayscale image, and the second boundary circle includes the boundaries of all second nasopharyngeal junctions in the second color image. The set of borders for all first nasopharyngeal junctions within the first boundary circle is defined as... ,in, This is the border of the area where the first nasopharynx connects to the larynx. Define the set of borders of all connected parts of the second nasopharynx within the second boundary circle as... ,in, This is the border of the area where the second nasopharynx connects. Then the registration components for each pair of nasopharyngeal junctions are: Based on the above parameters, the area crossover ratio of the first nasopharyngeal junction and the corresponding second nasopharyngeal junction is calculated, that is, the area crossover ratio of the registration components for each pair of nasopharyngeal junctions is calculated, and the formula is as follows: ;in, It represents the intersection of the areas where the first and second nasopharyngeal junctions meet. This is the union of the areas of the first and second nasopharyngeal junctions. The area intersection-union ratio of each pair of nasopharyngeal junction registration components can be calculated using the formula for the area intersection-union ratio of each pair of nasopharyngeal junction registration components, thus facilitating subsequent registration of the first and corresponding second nasopharyngeal junctions.

[0061] After calculating the area intersection-union ratio (IUGR) of each pair of nasopharyngeal junction registration components, the registration process of the nasopharyngeal junctions can be transformed into a task assignment problem of solving a single mapping under a minimum cost function. In the calculation of the IUGR, a larger IUGR value indicates a higher probability that the corresponding first and second nasopharyngeal junctions meet the registration requirements. Similarly, in solving the task assignment problem, a smaller cost value also indicates a higher probability that the first and corresponding second nasopharyngeal junctions can be registered. Specifically, the smaller the difference in the IUGR, the greater the probability that the first and corresponding second nasopharyngeal junctions meet the registration requirements. Based on this, in this application, the formula used to calculate the cost of the first and corresponding second nasopharyngeal junctions to be registered is as follows: ; in, This is the cost matrix. Specifically, by calculating the formula for registering the first nasopharyngeal junction and its corresponding second nasopharyngeal junction, the cost of all first nasopharyngeal junctions and their corresponding second nasopharyngeal junctions can be calculated during the registration process. Based on this cost calculation result, the second nasopharyngeal junction that matches the first nasopharyngeal junction and has the largest area intersection-union ratio can be selected. This registration result will provide valuable support for subsequently obtaining the same and different features in the third grayscale image and the third color image.

[0062] like Figure 2 As shown, step S3: Extract the common and different features of the nasopharyngeal region from the third grayscale image and the third color image.

[0063] like Figure 7 As shown, step S3: the step of extracting the common features of the nasopharyngeal region from the third grayscale image and the third color image includes: Step S31: Obtain the first image features of the nasopharyngeal junction in the third grayscale image and the second image features of the nasopharyngeal junction in the third color image. In this application, the first image features specifically refer to the image features of the nasopharyngeal junction in the third grayscale image, and the second image features specifically refer to the image features of the nasopharyngeal junction in the third color image. By performing corresponding processing operations on the first and second image features, the necessary foundation can be provided for subsequent detection of nasopharyngeal lesions, thereby achieving the detection target of nasopharyngeal lesions.

[0064] Step S32: Perform a cross-product on the first image features and the second image features. Element-wise cross-product includes element-wise multiplication between vectors and pattern-specific multiplication of tensors. In this application, element-wise cross-product is used to process the first image features and the second image features, multiplying the first image features and the second image features to obtain the same features required subsequently.

[0065] Step S33: Based on cross-product, enhance the similarities between the first and second image features in space and suppress the differences between the first and second image features. In this application, element-wise cross-product is used to multiply the first and second image features one by one. This operation enhances the image features of the areas where they are spatially activated together. At the same time, the values ​​of spatially differing image features tend to 0 due to this multiplication operation, thereby suppressing the differing image features and ultimately achieving the effect of suppressing the differing features. Through the above settings, the enhancement of similar features is essentially to enhance the image features that are spatially activated together in the third grayscale image and the third color image of the nasopharynx-laryngeal region. This enhancement can improve the recognition of such image features, providing strong support for subsequent analysis of the specific types of lesions in the nasopharynx-laryngeal region based on these image features, thereby further improving the accuracy of lesion detection in the nasopharynx-laryngeal region.

[0066] like Figure 8 As shown, step S3: Extract the differential features of the nasopharyngeal junction from the third grayscale image and the third color image. This includes: Step S3a: Obtain the first image features of the nasopharyngeal junction in the third grayscale image and the second image features of the nasopharyngeal junction in the third color image.

[0067] Step S3b: Subtract the first image feature from the second image feature to generate a first difference feature map. In this application, the third color image only has the ability to present the surface image features of the nasopharynx, while the third grayscale image can simultaneously present the surface and internal image features of the nasopharynx. Based on the difference in the presentation range of the two image features, the image feature portion in the first image feature that shares a common activation region with the second image feature can be removed through the above settings. After this processing, the first difference feature map can clearly present the image features that differ between the first image feature and the second image feature, providing clear evidence of difference features for subsequent related analysis.

[0068] For example, if there is internal inflammation or cancerous tendencies in the nasopharynx, it will appear as color blocks in the third grayscale image, but not in the third color image. Step S33 identifies internal inflammation or cancerous tendencies as a differential feature and suppresses its display. However, these differential features can facilitate the determination of whether lesions exist inside the nasopharynx. Therefore, step S3b can retain the portion of the third grayscale image containing color blocks as a differential feature, thereby facilitating subsequent determination of lesions in the nasopharynx based on similar and differential features.

[0069] Step S3c: Perform image processing on the first difference feature map to determine attention features. In the field of deep learning, attention features are a dynamic feature optimization mechanism. Their core function is to adaptively enhance the information representation of key feature channels while suppressing irrelevant redundant information. In this application, the image processing operation on the first difference feature map follows a specific process: first, the first difference feature map is sequentially processed by average pooling and max pooling; then, feature transformation is performed through a shared multilayer perceptron; next, channel addition is performed; and finally, feature output is completed through an activation function. Through the above series of steps, corresponding attention features can be generated. These attention features can effectively enhance the target image features in the first difference feature map, providing necessary support for subsequent accurate acquisition of difference features.

[0070] Step S3d: Multiply the attention feature and the first difference feature map to generate a second difference feature map containing the difference features of the third grayscale image, thereby obtaining the difference features. This setup enhances the display of difference features, which in turn facilitates the acquisition of difference features from the second difference feature map.

[0071] like Figure 2 As shown, step S4 involves performing modality fusion processing on the extracted identical and differential features, merging the identical and differential features to generate a fusion map. Through the above settings, the fusion map can simultaneously present identical and differential features. Based on these two types of feature information provided by the fusion map, the specific type of lesions in the nasopharynx and larynx can be determined more efficiently, thus creating favorable conditions for improving the detection efficiency of lesions in the nasopharynx and larynx.

[0072] like Figure 9 As shown, generating the fused graph in step S4 includes: Step S41: Combine identical and dissimilar features, and align them to generate a fusion image. This involves aligning and merging images containing identical features with a second dissimilar feature image. The fusion image allows comparison of the presence of identical or dissimilar graphic features at the same location in the nasopharynx, enabling subsequent lesion detection in the nasopharynx. The specific formula for generating the fusion image is as follows: ; in, To generate the fusion graph, For feature maps containing the same features, This is the second difference feature map containing the difference features. It is a 1×1 convolution. That is, the formula for generating the fusion map can be used to concatenate similar and dissimilar features to generate the fusion map.

[0073] like Figure 2 As shown, step S5: Based on the generated fusion image, and combined with the extracted similar and different features, the lesion type and cause of the lesion in the nasopharynx are determined, thereby realizing the detection of lesions in the nasopharynx.

[0074] In this application, when a lesion occurs in the nasopharynx, the area corresponding to the nasopharynx in the third grayscale image or the third color image will exhibit image features related to the lesion. Specifically, the third grayscale image has the ability to simultaneously present both surface and internal image features of the nasopharynx, while the third color image can only present image features present on the surface of the nasopharynx. Based on the difference in the range of image features presented by the two, it can be inferred that: image features common to both the third grayscale image and the third color image correspond to lesions on the surface of the nasopharynx, that is, by analyzing the same features, the lesions on the surface of the nasopharynx can be determined; while image features present only in the third grayscale image and not in the third color image correspond to lesions inside the nasopharynx, that is, by analyzing the different features, the lesions inside the nasopharynx can be determined.

[0075] In summary, the nasopharyngeal disease image recognition method acquires a first grayscale image and a first color image containing the nasopharyngeal region, performs image cross-correlation on the first grayscale image and the first color image to generate a second grayscale image and a second color image, then registers the nasopharyngeal connecting parts in the second grayscale image and the second color image to generate a third grayscale image and a third color image. Finally, based on the difference and similarity features of the nasopharyngeal connecting parts in the third grayscale image and the third color image, lesions in the nasopharyngeal region are detected, thereby achieving joint diagnosis of lesion types in the nasopharyngeal region and correlation analysis of lesion causes, which helps to improve the accuracy of lesion detection methods in detecting nasopharyngeal lesion types.

[0076] In this application, the neural network model, leveraging convolutional normalization and adaptive convolution techniques, performs multiple extraction operations from shallow to deep layers on the image features of the third color image and the third grayscale image. Specifically, it processes the image through attention convolutional layers and standard convolutional modules. Shallow feature extraction captures shape features such as texture, color, and edges in the third color and grayscale images; deep feature extraction determines the specific category of features in the third color and grayscale images, such as identifying inflammation or surface wounds. The neural network model extracts modal features from the third color and grayscale images separately through the stacked structure of the aforementioned convolutional layers. The extracted features are then fused using a multimodal feature fusion module to obtain both identical and differing features. Based on identical features, it can detect surface lesions in the nasopharynx, such as inflammation caused by wounds or local masses caused by tumors; based on differing features, it can detect nodules present within the nasopharynx. After lesion detection, the identical and differing features are stitched together and their dimensions aligned through channel stitching operations to obtain a fused image. The fused image is then fed into subsequent attention-based convolutional layers for further feature extraction. This allows the neural network model to grasp the causal relationship between different image features and lesion types through deep learning, thereby enabling the diagnosis of different lesion types and their causes in the nasopharynx and larynx. This process helps the neural network model conduct joint diagnosis of lesion types and correlation analysis of lesion causes, ultimately improving its accuracy in detecting lesion types.

[0077] like Figure 10 As shown, this embodiment also provides a nasopharyngeal disease image recognition lesion system, which includes an acquisition device 21 and a nasopharyngeal disease image recognition lesion device. The acquisition device 21 is used to acquire a first grayscale image and a first color image, and the nasopharyngeal disease image recognition lesion device is used to perform a nasopharyngeal disease image recognition method.

[0078] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. In the above embodiments of the present invention, the descriptions of each embodiment have their own emphasis; parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. It should be understood that the disclosed technical content in the several embodiments provided in this application can be implemented in other ways.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An AI-based image recognition method for nasopharyngeal lesions, characterized in that, Includes the following steps: A first grayscale image and a first color image containing the nasopharynx are acquired, and the regions containing the nasopharynx in the first grayscale image and the regions containing the nasopharynx in the first color image are cross-correlated to generate a second grayscale image and a second color image; the regions containing the nasopharynx in the second grayscale image and the regions containing the nasopharynx in the second color image are roughly aligned in space. For the nasopharyngeal junction in the second grayscale image, a registration process is performed with the nasopharyngeal junction in the second color image. Through the registration process, it is ensured that the nasopharyngeal junctions with the same spatial position in the second grayscale image and the second color image form a corresponding relationship, and a third grayscale image and a third color image are generated; the spatial position of the same nasopharyngeal junction in the third grayscale image and the third color image is completely consistent. Extract the common and different features of the nasopharyngeal junction from the third grayscale image and the third color image; Modality fusion processing is performed on the extracted identical features and differential features to merge the identical features and differential features, thereby generating a fusion map; Based on the fusion diagram and the similar and different features, the lesion sites in the nasopharynx are identified.

2. The AI-based image recognition method for nasopharyngeal lesions according to claim 1, characterized in that, The device that acquires the first grayscale image is defined as a grayscale device, and the device that acquires the first color image is defined as a color image device. The image cross-correlation includes: The height, pixel pitch, and focal length of the grayscale device and the color image device are obtained, and the width and height data of the first color image and the first grayscale image are obtained simultaneously. Based on the height, pixel pitch, and focal length of the acquired grayscale and color image devices, the resolution quotient of the grayscale and color image devices is calculated. Based on the calculated resolution quotient and the width and height information of the acquired first color image and first grayscale image, the first resampling factor corresponding to the first color image is determined; A second resampling factor is determined based on the first resampling factor, and the first color image is resampled using the second resampling factor to generate a resampled color image. Image cross-correlation processing is performed on the region where the nasopharynx is located in the first grayscale image and the region where the nasopharynx is located in the resampled color image.

3. The AI-based image recognition method for nasopharyngeal lesions according to claim 1, characterized in that, The nasopharyngeal connection in the second grayscale image is defined as the first nasopharyngeal connection, and the nasopharyngeal connection in the second color image is defined as the second nasopharyngeal connection. The registration of the nasopharyngeal junction in the second grayscale image and the nasopharyngeal junction in the second color image includes the following steps: Obtain the spatial location information of the first nasopharyngeal junction and the second nasopharyngeal junction; Based on the obtained spatial location, the intersection area and the union area between the first nasopharyngeal connection point and its corresponding second nasopharyngeal connection point are calculated respectively. Then, the area intersection-union ratio between the first nasopharyngeal connection point and the corresponding second nasopharyngeal connection point is calculated using the intersection area and the union area. Filter and remove the first nasopharyngeal junction and its corresponding second nasopharyngeal junction with an area intersection ratio less than a preset threshold.

4. The AI-based image recognition method for nasopharyngeal lesions according to claim 1, characterized in that, Obtaining the same features of the nasopharyngeal junction in the third grayscale image and the third color image includes the following steps: Obtain the first image feature corresponding to the nasopharyngeal connection in the third grayscale image, and the second image feature corresponding to the nasopharyngeal connection in the third color image; Perform a cross-product operation on the obtained first image features and second image features.

5. A method for image recognition of nasopharyngeal lesions based on AI according to claim 1 or 4, characterized in that, The steps to obtain the differential features of the nasopharyngeal junction in the third grayscale image and the third color image include: Obtain the first image features of the nasopharynx-laryngeal junction in the third grayscale image, and the second image features of the nasopharynx-laryngeal junction in the third color image; The first image feature is subtracted from the second image feature, and the first difference feature map is generated by this operation; Image processing operations are performed on the generated first difference feature map, and attention features are determined based on the processing results; The product of the acquired attention features and the first image features is calculated, and a second difference feature map is generated by the product. The second difference feature map contains the difference features in the third grayscale image features, thereby realizing the acquisition of the difference features.

6. The AI-based image recognition method for nasopharyngeal lesions according to claim 5, characterized in that, The generation of the fusion graph includes the following steps: The same and different features are concatenated. Based on a single convolution operation, the concatenated same and different features are aligned to generate a fused image.

7. The AI-based image recognition method for nasopharyngeal lesions according to claim 6, characterized in that, The model for generating the fusion graph is defined as a neural network model, and the generation of the fusion graph includes: The third grayscale image and the third color image are input into the neural network model; Under the action of the neural network model, the third grayscale image and the third color image are sequentially subjected to operations such as convolution normalization, feature extraction, adaptive convolution, structural reparameterization, fusion of data residual structures, and feature concatenation, thereby generating the same features and different features through these operations. The obtained identical feature maps and differing features are spliced ​​together to generate a fused map.

8. The AI-based image recognition method for nasopharyngeal lesions according to claim 2, characterized in that, The step of performing image cross-correlation on the nasopharyngeal region in the first grayscale image and the nasopharyngeal region in the resampled color image also includes the following steps: A first boundary circle is set along the contour of the nasopharynx in the first grayscale image; at the same time, a second boundary circle is set along the contour of the nasopharynx in the resampled color image. The first center coordinates of the nasopharynx in the first grayscale image, the second center coordinates of the nasopharynx in the resampled color image, the width and height of the first boundary circle, and the width and height of the second boundary circle are obtained, and the relative positions of the nasopharynx in the first grayscale image and the resampled color image are matched accordingly. Based on the relative positions of the nasopharynx and larynx in the first grayscale image and the resampled color image, image cross-correlation processing is performed on the regions where the nasopharynx and larynx are located in the first grayscale image and the regions where the nasopharynx and larynx are located in the resampled color image.

9. The AI-based image recognition method for nasopharyngeal lesions according to claim 8, characterized in that, The nasopharyngeal junction in the second grayscale image is defined as the first nasopharyngeal junction, and the nasopharyngeal junction in the second color image is defined as the second nasopharyngeal junction. A third boundary circle is set along the contour of the first nasopharyngeal junction, and a fourth boundary circle is set along the contour of the second nasopharyngeal junction. The registration of the nasopharyngeal junction in the second grayscale image and the color image also includes the following steps: Based on the area intersection-union ratio of the first nasopharyngeal connection and the second nasopharyngeal connection, the first center coordinate, the second center coordinate, the width and height of the third boundary circle, and the width and height of the fourth boundary circle, a registration operation is performed on the first nasopharyngeal connection and the second nasopharyngeal connection, which are spatially closest in the second grayscale image and the second color image, and a nasopharyngeal connection registration component is generated through this registration. The registration components of the nasopharyngeal connection area with an area intersection ratio less than a preset threshold are filtered out and removed.

10. An AI-based image recognition system for nasopharyngeal lesions, characterized in that, include: The acquisition device is used to acquire the first grayscale image and the first color image as described in any one of claims 1 to 9; A computer device for performing the image recognition method for nasopharyngeal diseases as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Nasopharyngeal-carcinoma (NPC) lesion automatic-segmentation method and nasopharyngeal-carcinoma lesion automatic-segmentation systems based on deep learning

    CN108257134A

  • Multi-modal nasopharyngeal carcinoma image segmentation method based on attention and graph convolution

    CN115187613A

  • Image segmentation method, electronic equipment and storage medium

    CN117058164A

  • Nasopharynx cancer pseudo-CT generation method based on multi-modal deformation registration and Transform network

    CN119625110A

  • Defect detection method and detection equipment for photovoltaic module

    CN120782761A