ENT Intelligent Diagnosis and Monitoring System Based on Multimodal Data Fusion
The endoscopic monitoring images are processed through edge detection and pixel decision window technology, and the correlation map is generated in combination with monitoring text data, which solves the problem of low image quality of nasopharyngeal endoscopics, and achieves efficient fusion of multimodal data and improves diagnostic accuracy.
Patent Information
- Application Number
- CN202510605301.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-12
AI Technical Summary
In the prior art, the image quality of the nasopharyngeal endoscopy monitors is affected by perturbations such as uneven light, reflection, bright mucus occlusion and dynamic blur, and cannot effectively realize the cross-modal alignment and unified semantic expression of multimodal data, affecting the accuracy of diagnosis of ENT diseases.
The edges of the endoscopic monitoring image were extracted using the Canny edge detection algorithm, and secretion perturbation mask was generated. The restored image was constructed through XOR operation, and the pixel decision window was adaptively set for each pixel point for information decision-making, and the feature fusion was used to generate an association map.
It significantly improves the structural clarity and lesion visibility of the endoscopic monitoring images, suppresses the noise caused by secretions and fuzzy occlusion, realizes information complementarity and semantic enhancement of multimodal data, and improves the knowledge fusion ability of ENT monitoring.
Smart Images

Figure CN120107271B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of multimodal data fusion. More specifically, this application relates to an intelligent diagnosis and monitoring system for otolaryngology based on multimodal data fusion. Background Art
[0002] The intelligent diagnosis and monitoring of otolaryngology based on multimodal data fusion aims to build a comprehensive, accurate, and efficient intelligent medical system by integrating multi-source heterogeneous data such as nasopharyngeal endoscopy images, voice audio, physiological monitoring signals, and structured medical record texts, so as to achieve the auxiliary diagnosis of common otolaryngology diseases and early lesions, the identification of the evolution trend of the disease course, and continuous health monitoring. This method not only breaks through the limitations of single-modal information expression but also enhances the understanding of complex lesion morphology, symptom interaction, and spatio-temporal evolution process by deeply combining artificial intelligence and medical knowledge. The overall process of the system includes multimodal data acquisition, inter-modal feature alignment and semantic fusion, disease recognition and severity assessment driven by fused features, health status monitoring based on atlases and time-series models, and a visualization feedback and intelligent interaction platform for doctors and patients. Especially in the diagnosis and treatment of diseases in the nasopharynx, throat and other parts, multimodal collaboration has natural advantages: images can directly present the lesion morphology, such as swelling, inflammation, masses, etc.; voice data indirectly reflects the functional status of the vocal cords or throat, such as hoarseness, abnormal pitch, etc.; physiological signal records reflect potential systemic symptoms such as sleep apnea and heart rate, while text medical records supplement unstructured knowledge such as medical history and drug use.
[0003] However, in the prior art, various modalities are distributed differently in the physical space and semantic space, and it is impossible to achieve effective cross-modal alignment and construct a unified semantic expression space. Moreover, there are disturbances such as uneven illumination, reflection, high-brightness mucus occlusion, and motion blur in the endoscopic monitoring images of the nasopharynx, as well as non-structured sound interferences such as coughing and swallowing during the patient's pronunciation process, which seriously affect the data quality. Therefore, how to improve the image quality of endoscopic monitoring images and generate an associated atlas of patients in combination with monitoring texts to achieve knowledge fusion in otolaryngology monitoring is a difficult problem faced by the industry. Summary of the Invention
[0004] This application provides an intelligent diagnosis and monitoring system for otolaryngology based on multimodal data fusion, which can improve the image quality of endoscopic monitoring images and generate an associated atlas of patients in combination with monitoring texts to achieve knowledge fusion in otolaryngology monitoring.
[0005] This application provides an intelligent diagnosis and monitoring system for otolaryngology based on multimodal data fusion, and the diagnosis and monitoring system includes:
[0006] An image acquisition module, configured to acquire multimodal monitoring data of a patient during otolaryngology monitoring, and extract an endoscopic monitoring image of the patient from the multimodal monitoring data;
[0007] An image restoration module, configured to perform edge monitoring on the endoscopic monitoring image to obtain an endoscopic monitoring edge image, generate a secretion perturbation mask based on the endoscopic monitoring edge image, and determine an endoscopic monitoring restoration image of the patient according to the secretion perturbation mask and the endoscopic monitoring image;
[0008] An information decision module, configured to determine a pixel decision window for each pixel point in the endoscopic monitoring restoration image, perform information decision on each pixel point in the endoscopic monitoring restoration image based on the corresponding pixel decision window, and further obtain an endoscopic monitoring decision image of the patient;
[0009] An associated graph generation module, configured to extract monitoring text data of the patient from the multimodal monitoring data, perform feature fusion on the monitoring text data and the endoscopic monitoring decision image, and further generate a fusion associated graph of the patient during otolaryngology monitoring.
[0010] In this embodiment, the multimodal monitoring data includes an endoscopic monitoring image and monitoring text data of the patient during otolaryngology monitoring.
[0011] In this embodiment, the Canny edge monitoring algorithm is used to perform edge detection on the endoscopic monitoring image.
[0012] In this embodiment, generating a secretion perturbation mask based on the endoscopic monitoring edge image specifically includes:
[0013] Performing morphological processing on the endoscopic monitoring edge image to obtain a closed edge image;
[0014] Performing hole filling on the closed edge image to obtain a single-connected region image;
[0015] Performing erosion processing on the single-connected region image, and generating a secretion perturbation mask according to the erosion processing result.
[0016] In this embodiment, determining the endoscopic monitoring restoration image of the patient according to the secretion perturbation mask and the endoscopic monitoring image is to perform an exclusive OR operation on the secretion perturbation mask and the endoscopic monitoring image, and further obtain the endoscopic monitoring restoration image of the patient.
[0017] In this embodiment, determining a pixel decision window for each pixel point in the endoscopic monitoring restoration image specifically includes:
[0018] Obtaining a preset pixel window;
[0019] For each pixel point in the endoscopic monitoring restored image, determine the information fineness corresponding to the pixel point according to the pixel window;
[0020] Based on the information fineness, determine the pixel decision window of the pixel point, and then obtain the pixel decision windows of all pixel points in the endoscopic monitoring restored image.
[0021] In this embodiment, determining the information fineness corresponding to the pixel point according to the pixel window specifically includes:
[0022] Take the pixel point as the window center of the pixel window, and then determine all neighborhood windows of the pixel window;
[0023] Determine the pixel heterogeneity coefficient between the pixel window and each neighborhood window;
[0024] Determine the information fineness corresponding to the pixel point through all the pixel heterogeneity coefficients.
[0025] In this embodiment, determining the pixel decision window of the pixel point based on the information fineness specifically includes:
[0026] Obtain a preset fineness threshold;
[0027] When the information fineness is greater than the fineness threshold, set the size of the pixel decision window of the pixel point to 3 pixels × 3 pixels;
[0028] When the information fineness is not greater than the fineness threshold, set the size of the pixel decision window of the pixel point to 7 pixels × 7 pixels.
[0029] In this embodiment, performing information decision on each pixel point in the endoscopic monitoring restored image based on the corresponding pixel decision window, and then obtaining the endoscopic monitoring decision image of the patient specifically includes:
[0030] Determine the background background noise of the endoscopic monitoring restored image;
[0031] For each pixel point in the endoscopic monitoring restored image, determine the local information complexity and local pixel level of the pixel point according to the corresponding pixel decision window;
[0032] Perform information decision on the pixel point through the background background noise, the local information complexity, and the local pixel level to obtain the decision pixel point corresponding to the pixel point, and then obtain the decision pixel points corresponding to all pixel points in the endoscopic monitoring restored image;
[0033] Construct the endoscopic monitoring decision image of the patient based on all the decision pixel points.
[0034] In this embodiment, feature fusion is performed on the monitored text data and the endoscopic monitoring decision image, and then a fusion association map of the patient in the otolaryngology monitoring is generated, which specifically includes:
[0035] Perform semantic parsing on the monitored text data to obtain the structured knowledge of the patient in the otolaryngology monitoring;
[0036] Extract lesion feature vectors from the endoscopic monitoring decision image to obtain lesion image feature vectors;
[0037] Perform joint modeling on the structured knowledge and the lesion image feature vectors, and then generate a fusion association map of the patient in the otolaryngology monitoring.
[0038] The technical solution provided by the disclosed embodiment of the present application has the following beneficial effects:
[0039] The multi-modal monitoring data of the patient in the otolaryngology monitoring is obtained through the image acquisition module, and the endoscopic monitoring image of the patient is extracted from the multi-modal monitoring data; the image restoration module performs edge monitoring on the endoscopic monitoring image to obtain an endoscopic monitoring edge image, generates a secretion perturbation mask based on the endoscopic monitoring edge image, and determines the endoscopic monitoring restored image of the patient according to the secretion perturbation mask and the endoscopic monitoring image; the information decision module determines the pixel decision window of each pixel point in the endoscopic monitoring restored image, and performs information decision on each pixel point in the endoscopic monitoring restored image based on the corresponding pixel decision window, and then obtains the endoscopic monitoring decision image of the patient; the association map generation module extracts the monitored text data of the patient from the multi-modal monitoring data, performs feature fusion on the monitored text data and the endoscopic monitoring decision image, and then generates a fusion association map of the patient in the otolaryngology monitoring.
[0040] It can be seen that in this application, first, in the otorhinolaryngology endoscope monitoring, the processing flow of edge extraction for the endoscope monitoring image, generating a secretion perturbation mask, and constructing an endoscope monitoring restored image can effectively improve the structural clarity and lesion visibility of the image, and significantly suppress the image noise and information loss problems caused by secretions, blurred occlusion, etc.; then, for each pixel point in the endoscope monitoring restored image, a pixel decision window is adaptively determined, and local information decision is made accordingly, that is, a small window is used in the region with rich details to retain the edge and lesion contour, and a large window is used in the smooth region to suppress noise and artifacts, thereby significantly improving the overall image quality and structural clarity; finally, the monitoring text data and the endoscope monitoring decision image are feature-fused to generate a fusion correlation map of the patient in the otorhinolaryngology monitoring. By fusing the structured semantic information in the text and the lesion region features in the image, the information complementarity and semantic enhancement between multiple modalities can be realized, the image understanding ability and semantic clarity can be improved, the interference regions such as secretions are removed through image decision optimization processing, and the overall image quality can be improved, so as to achieve knowledge fusion in otorhinolaryngology monitoring.
[0041] In summary, the technical solution adopted in this application can improve the image quality of the endoscope monitoring image, and generate an association map of the patient in combination with the monitoring text to achieve knowledge fusion in otorhinolaryngology monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0043] Figure 1 is a module structure diagram of an otorhinolaryngology intelligent diagnosis and monitoring system provided by the present application;
[0044] Figure 2 is an exemplary flowchart for generating a secretion perturbation mask provided by the present application;
[0045] Figure 3 is an exemplary flowchart for determining the pixel decision window of each pixel point in the endoscope monitoring restored image provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0047] The embodiment of the present application provides an otolaryngology intelligent diagnosis and monitoring system based on multimodal data fusion. Its core is to obtain multimodal monitoring data of a patient during otolaryngology monitoring through an image acquisition module, and extract the endoscopic monitoring image of the patient from the multimodal monitoring data; an image restoration module performs edge monitoring on the endoscopic monitoring image to obtain an endoscopic monitoring edge image, generates a secretion perturbation mask based on the endoscopic monitoring edge image, and determines the endoscopic monitoring restored image of the patient according to the secretion perturbation mask and the endoscopic monitoring image; an information decision module determines the pixel decision window of each pixel point in the endoscopic monitoring restored image, performs information decision on each pixel point in the endoscopic monitoring restored image based on the corresponding pixel decision window, and then obtains the endoscopic monitoring decision image of the patient; an associated graph generation module extracts the monitoring text data of the patient from the multimodal monitoring data, performs feature fusion on the monitoring text data and the endoscopic monitoring decision image, and then generates an integrated associated graph of the patient during otolaryngology monitoring. Adopting the above solution can improve the image quality of the endoscopic monitoring image and generate an associated graph of the patient in combination with the monitoring text to achieve knowledge fusion in otolaryngology monitoring.
[0048] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners. Refer to Figure 1 As shown in the figure, which is a module structure diagram of an otolaryngology intelligent diagnosis and monitoring system based on multimodal data fusion according to an embodiment of the present application. The diagnosis and monitoring system includes: an image acquisition module 100, an image restoration module 200, an information decision module 300, and an associated graph generation module 400, which are described as follows:
[0049] The image acquisition module 100 is used to obtain multimodal monitoring data of a patient during otolaryngology monitoring and extract the endoscopic monitoring image of the patient from the multimodal monitoring data.
[0050] In specific implementation, first, the multi-modal monitoring data of the patient's ear, nose, and throat can be collected through various channels and devices. In this application, the multi-modal monitoring data includes the endoscopic monitoring images and monitoring text data during the ear, nose, and throat monitoring of the patient. The endoscopic monitoring images can be collected by a high-definition electronic nasopharyngeal and laryngoscope. The endoscopic monitoring images contain anatomical structures such as the nasopharynx, throat, and glottis, often accompanied by disturbances such as mucus and reflection. The monitoring text data includes the text format of voice audio data, physiological monitoring data, and structured / unstructured medical record information. In actual implementation, the voice audio data can be collected by a high-sensitivity microphone and a recording module attached to the endoscope, the physiological monitoring data can be collected by a portable sleep monitor or a smart vital sign acquisition patch, and the structured / unstructured medical record information of the patient can be obtained through the hospital HIS system; then, the endoscopic monitoring images of the patient can be extracted from the multi-modal monitoring data through data traversal and retrieval methods.
[0051] The image restoration module 200 is used to perform edge monitoring on the endoscopic monitoring image to obtain an endoscopic monitoring edge image, generate a secretion disturbance mask based on the endoscopic monitoring edge image, and determine the endoscopic monitoring restored image of the patient according to the secretion disturbance mask and the endoscopic monitoring image.
[0052] In specific implementation, the Canny edge monitoring algorithm can be used to perform edge detection on the endoscopic monitoring image to obtain an endoscopic monitoring edge image; it should be noted that the Canny edge detection algorithm efficiently detects the edges of the endoscopic monitoring image through multi-stage processing (such as noise suppression, gradient calculation, non-maximum suppression, etc.), so as to obtain the endoscopic monitoring edge image.
[0053] Preferably, in this embodiment, refer to Figure 2 As shown, this figure is an exemplary flowchart for generating a secretion disturbance mask in an embodiment of this application. In this embodiment, generating a secretion disturbance mask based on the endoscopic monitoring edge image can be specifically implemented by the following steps:
[0054] In step S21, morphological processing is performed on the endoscopic monitoring edge image to obtain a closed edge image;
[0055] In step S22, hole filling is performed on the closed edge image to obtain a single-connected region image;
[0056] In step S23, erosion processing is performed on the single-connected region image, and a secretion disturbance mask is generated based on the erosion processing result.
[0057] In specific implementation, first, morphological processing can be performed on the endoscopic monitoring edge image, that is, the closing operation is used to process the endoscopic monitoring edge image. The dilation operation is used to fill small gaps in the edge, and then the erosion operation is used to restore the original shape of the edge. This operation can eliminate the problems of edge breakage or discontinuity, making the edges of the image background area coherent and closed, so that a closed edge image can be obtained. Among them, the structuring element used in the closing operation can be selected as circular or elliptical to adapt to the characteristics of the edges of the background area in the endoscopic monitoring image. Then, hole filling can be performed on the closed edge image. It should be noted that the purpose of hole filling is to completely fill the pixels inside the closed area with the foreground value and eliminate the holes inside the edge, thereby forming a complete and continuous single-connected area, that is, a single-connected area image. Furthermore, erosion processing can be performed on the single-connected area image, that is, for the filled single-connected area image, the morphological erosion operation is performed to remove the irregular or adhered parts of the edge and narrow the boundary of the secretion area. The erosion operation can further eliminate the artifacts or abnormal edges in the non-secretion area, making the mask of the secretion area more accurate. Finally, a secretion perturbation mask can be generated based on the result of the erosion processing, that is, according to the result of the erosion processing, the eroded area is extracted as the secretion perturbation mask. Among them, the secretion perturbation mask is a mask image used to remove the perturbation of the secretion area, and the pixel values corresponding to the secretion perturbation area in the secretion perturbation mask are set to 1, and the pixel values of the remaining areas are set to 0.
[0058] In this embodiment, determining the endoscopic monitoring restoration image of the patient according to the secretion perturbation mask and the endoscopic monitoring image is to perform an exclusive OR operation on the secretion perturbation mask and the endoscopic monitoring image, and then obtain the endoscopic monitoring restoration image of the patient.
[0059] In specific implementation, first, the secretion perturbation mask and the original endoscopic monitoring image are pixel-level corresponding; then, the exclusive OR operation can be performed on each pixel in the secretion perturbation mask and the corresponding pixel in the endoscopic monitoring image. The logical rule is that when the pixel value in the secretion perturbation mask is 1 (indicating the secretion perturbation area), the exclusive OR operation is performed with the corresponding pixel in the endoscopic monitoring image, and the result is 0, that is, the perturbation information in this area is eliminated or cleared; when the mask pixel value is 0 (indicating the non-secretion perturbation area), the exclusive OR operation is performed with the corresponding pixel in the endoscopic monitoring image, and the result retains the original pixel value of the endoscopic monitoring image. It should be noted that the result of the exclusive OR operation is the endoscopic monitoring restoration image of the patient. Performing the exclusive OR operation on the secretion perturbation mask and the original image can clear or weaken the interference factors without damaging the normal structure and improve the image diagnosis value.
[0060] It should be noted that in the monitoring of ENT endoscopes, the processing flow of edge extraction, generation of secretion perturbation masks, and construction of restored endoscope monitoring images for endoscope monitoring images can effectively improve the structural clarity and lesion visibility of the images, significantly suppress image noise and information loss caused by secretions, blurring, occlusion, etc., not only enhance the edge contrast of key anatomical regions (such as nasal cavities, pharyngeal walls, mucosal folds, etc.), but also provide a more stable basic image for subsequent image segmentation and lesion detection.
[0061] The information decision module 300 is used to determine the pixel decision window of each pixel point in the restored endoscope monitoring image, make information decisions on each pixel point in the restored endoscope monitoring image based on the corresponding pixel decision window, and then obtain the endoscope monitoring decision image of the patient.
[0062] Preferably, in this embodiment, referring to Figure 3 As shown, this figure is an exemplary flowchart for determining the pixel decision window of each pixel point in the restored endoscope monitoring image in the embodiment of the present application. In this embodiment, the pixel decision window of each pixel point in the restored endoscope monitoring image can be specifically implemented by the following steps:
[0063] In step S31, obtain a pre-set pixel window;
[0064] In step S32, for each pixel point in the restored endoscope monitoring image, determine the information fineness corresponding to the pixel point according to the pixel window;
[0065] In step S33, determine the pixel decision window of the pixel point based on the information fineness, and then obtain the pixel decision window of each pixel point in the restored endoscope monitoring image.
[0066] When specifically implemented, first, the size of the pixel window can be pre-set according to historical experience and data analysis, which will not be elaborated here; then, for each pixel point in the restored endoscope monitoring image, the information fineness corresponding to the pixel point can be determined according to the pixel window, so as to determine the pixel decision window of the pixel point based on the information fineness. Among them, the pixel decision window is a pixel window centered on the corresponding pixel point. Through the above method, the pixel decision window of each pixel point in the restored endoscope monitoring image can be obtained.
[0067] In this embodiment, the information fineness corresponding to the pixel point can be specifically determined according to the pixel window in the following way, that is:
[0068] Take the pixel point as the window center of the pixel window, and then determine all the neighborhood windows of the pixel window;
[0069] Determine the pixel heterogeneity coefficient between the pixel window and each neighborhood window;
[0070] Determine the information fineness corresponding to the pixel point based on all the pixel heterogeneity coefficients.
[0071] In specific implementation, first, a pixel point can be used as the window center of a pixel window, and then multiple neighborhood windows of the same size as the pixel window can be set around the current pixel window; then, the pixel heterogeneity coefficients between the pixel window and each neighborhood window can be determined, where the pixel heterogeneity coefficient represents the degree of difference between the pixels included in the pixel window and the pixels included in the neighborhood window. In actual implementation, the mean square error of the pixel values between the pixel point included in the pixel window and the pixel point included in the neighborhood window can be calculated, and the calculation result can be used as the pixel heterogeneity coefficient between the pixel window and the neighborhood window. Through the above method, the pixel heterogeneity coefficients between the pixel window and each neighborhood window can be obtained; finally, the information fineness corresponding to the pixel point can be determined based on all the pixel heterogeneity coefficients, where the information fineness represents the information fineness degree of the area corresponding to the pixel point. The more detailed information is included in the area corresponding to the pixel, the greater the difference between the area corresponding to the pixel point and other areas. In actual implementation, the mean value of all the pixel heterogeneity coefficients can be used as the information fineness corresponding to the pixel point.
[0072] In this embodiment, the pixel decision window of the pixel point can be determined based on the information fineness by specifically adopting the following method, that is:
[0073] Obtain a preset fineness threshold.
[0074] When the information fineness is greater than the fineness threshold, set the size of the pixel decision window of the pixel point to 3 pixels × 3 pixels.
[0075] When the information fineness is not greater than the fineness threshold, set the size of the pixel decision window of the pixel point to 7 pixels × 7 pixels.
[0076] In specific implementation, first, the fineness threshold can be preset based on experience or experiments according to different endoscopic monitoring image features or specific application scenarios, which will not be elaborated here; then, when the information fineness is greater than the fineness threshold, it indicates that the information fineness of the area where the pixel point is located is relatively large, meaning that the organizational structure of this area is complex and the texture changes are rich. In order to enhance the ability to preserve local features, a smaller 3×3 pixel window is used for information decision processing to ensure accuracy, that is, the size of the pixel decision window of the pixel point can be set to 3 pixels × 3 pixels; when the information fineness is not greater than the fineness threshold, it indicates that the area where the pixel point is located is relatively uniform and has little change. A larger 7×7 pixel window can be used for information decision, which helps to reduce noise interference and improve stability, that is, the size of the pixel decision window of the pixel point can be set to 7 pixels × 7 pixels.
[0077] In this embodiment, information decision-making is performed on each pixel point in the endoscopic monitoring restored image based on the corresponding pixel decision window, and then the endoscopic monitoring decision image of the patient can be obtained in the following specific manner, that is:
[0078] Determine the pathological background noise of the endoscopic monitoring restored image;
[0079] For each pixel point in the endoscopic monitoring restored image, determine the local information complexity and local pixel level of the pixel point according to the corresponding pixel decision window;
[0080] Perform information decision-making on the pixel point through the pathological background noise, the local information complexity, and the local pixel level to obtain the decision pixel point corresponding to the pixel point, and then obtain the decision pixel points corresponding to each pixel point in the endoscopic monitoring restored image;
[0081] Construct the endoscopic monitoring decision image of the patient based on all the decision pixel points.
[0082] Specifically, first, the pathological background noise of the endoscopic monitoring restored image can be determined. The pathological background noise represents the background noise level in the endoscopic monitoring restored image. The background noise variance of the endoscopic monitoring restored image can be calculated, and the calculation result can be used as the pathological background noise of the endoscopic monitoring restored image. Then, for each pixel point in the endoscopic monitoring restored image, the local information complexity and local pixel level of the pixel point can be determined according to the corresponding pixel decision window. The local information complexity represents the complexity of the pixel information contained in the pixel decision window corresponding to the pixel point, and the variance of all pixel points contained in the pixel decision window corresponding to the pixel point can be used as the local information complexity of the pixel point. The local pixel level represents the average pixel value contained in the pixel decision window corresponding to the pixel point.
[0083] In addition, specifically, information decision-making can be performed on the pixel point through the pathological background noise, the local information complexity, and the local pixel level to obtain the decision pixel point corresponding to the pixel point. In actual implementation, the pixel value of the decision pixel point corresponding to the pixel point can be determined by the following formula:
[0084]
[0085] Among them, represents the pixel value of the decision pixel point corresponding to pixel point i, represents the pixel value of pixel point i, represents the local information complexity of pixel point i, represents the local pixel level of pixel point i, It represents the pathological background noise. Through the above method, the decision pixel points corresponding to each pixel point in the endoscopic monitoring restored image can be obtained. Finally, the endoscopic monitoring decision image of the patient can be constructed based on all the decision pixel points, that is, all the decision pixel points are combined according to the corresponding image coordinates to obtain the endoscopic monitoring decision image of the patient.
[0086] It should be noted that adaptively determining the pixel decision window for each pixel point in the endoscopic monitoring restored image and making local information decisions accordingly helps to finely adjust the image processing strategy: small windows are used in areas with rich details to preserve edges and lesion contours, and large windows are used in smooth areas to suppress noise and artifacts, thus significantly improving the overall image quality and structural clarity. By enhancing the edge distinguishability and regional consistency of the image, it can effectively assist subsequent image semantic segmentation, anomaly recognition, and tissue structure determination.
[0087] The association graph generation module 400 is used to extract the monitoring text data of the patient from the multi-modal monitoring data, perform feature fusion on the monitoring text data and the endoscopic monitoring decision image, and then generate the fusion association graph of the patient in the otolaryngology monitoring.
[0088] Specifically, the monitoring text data of the patient can be extracted from the multi-modal monitoring data by traversing and retrieving.
[0089] In this embodiment, the following method can be specifically adopted to perform feature fusion on the monitoring text data and the endoscopic monitoring decision image and then generate the fusion association graph of the patient in the otolaryngology monitoring, that is:
[0090] Perform semantic parsing on the monitoring text data to obtain the structured knowledge of the patient in the otolaryngology monitoring;
[0091] Extract the lesion feature vectors from the endoscopic monitoring decision image to obtain the lesion image feature vectors;
[0092] Perform joint modeling on the structured knowledge and the lesion image feature vectors, and then generate the fusion association graph of the patient in the otolaryngology monitoring.
[0093] In specific implementation, first, semantic parsing can be performed on the monitored text data, that is, using natural language processing techniques to process the monitored text data, including steps such as word segmentation, named entity recognition (NER), and relation extraction. Through these techniques, key information in the text is extracted, so as to convert the extracted text information into structured data, that is, the structured knowledge of the patient in otolaryngology monitoring. For example, preprocessing text data such as otolaryngology-related examination records, doctor's notes, and postoperative summaries, including removing noise characters and unifying medical terms, using a pre-trained language model in the medical field, combined with a custom otolaryngology disease entity dictionary, to identify key entities in the text, such as lesion names, anatomical locations, lesion properties, and time status, etc., and using dependency syntactic analysis, attention mechanism or relation extraction model to identify the semantic relationships between entities, forming structured knowledge triples as the semantic knowledge representation of the patient in otolaryngology monitoring; then, lesion feature extraction can be performed on the endoscopic monitoring decision image, that is, through a deep learning model (such as a convolutional neural network CNN) to automatically identify and segment inflammatory patches, secretion-covered areas, cyst contours, etc. in the endoscopic monitoring decision image, so as to extract various features of the endoscopic monitoring decision image area, such as texture features (gray-level co-occurrence matrix, local binary pattern, etc.), morphological features (area size, shape, boundary, etc.), color features, edge features, etc., and use the vector composed of all the extracted features as the lesion image feature vector; finally, joint modeling can be performed on the structured knowledge and the lesion image feature vector, that is, joint modeling techniques can be used to fuse the structured knowledge and the lesion image feature vector to form a joint representation of multimodal data. For example, a multimodal neural network in deep learning (for example, a cross-modal learning model based on Transformer) considers the structured knowledge and the lesion image feature vector information at the same time. On the basis of joint modeling, graph construction techniques are used to convert the structured knowledge and the lesion image feature vector into nodes in the knowledge graph, and the relationships between the nodes are described as edges. Through the above method, the structured knowledge and the lesion image feature vector can be connected to form a fusion association graph of the patient in otolaryngology monitoring.
[0094] It should be noted that fusing the features of the monitored text data and the endoscopic monitoring decision image to generate a fusion association graph of the patient in otolaryngology monitoring has significant clinical and intelligent analysis value. By fusing the structured semantic information in the text and the lesion area features in the image, information complementarity and semantic enhancement between multimodals are achieved, which can improve the image understanding ability and semantic clarity. By optimizing the image decision-making process to remove interference areas such as secretions, the overall image quality is improved. The generated fusion association graph constructs an individualized knowledge network for the patient, facilitating time series analysis, pathological evolution reasoning, and intelligent assisted diagnosis.
[0095] By extracting the monitoring text data of patients from the pathological information and generating the fusion correlation map of patients in combination with the endoscopic monitoring decision-making images, the structured text data can be deeply fused with the image features, providing comprehensive patient pathological information. By correlating the image features of the lesion area with the clinical diagnosis and staging information in the report, the interpretation accuracy and visualization effect of the images can be enhanced, and the automated integration of pathological knowledge can also be achieved.
[0096] It can be seen that in this application, first, in the otolaryngology endoscopic monitoring, the processing flow of edge extraction, generation of secretion perturbation mask and construction of endoscopic monitoring restored image for the endoscopic monitoring images can effectively improve the structural clarity and lesion visibility of the images, and significantly suppress the image noise and information loss problems caused by secretions, blurring and occlusion, etc. Then, for each pixel point in the endoscopic monitoring restored image, an adaptive pixel decision window is determined, and local information decision is carried out accordingly, that is, a small window is used in the area with rich details to retain the edge and lesion contour, and a large window is used in the smooth area to suppress noise and artifacts, thus significantly improving the overall image quality and structural clarity. Finally, the monitoring text data and the endoscopic monitoring decision-making images are feature-fused to generate the fusion correlation map of patients in otolaryngology monitoring. By fusing the structured semantic information in the text with the lesion area features in the image, the information complementarity and semantic enhancement between multi-modalities can be realized, the understanding ability and semantic clarity of the images can be improved, the interference areas such as secretions are removed through image decision optimization processing, the overall image quality is improved, and the knowledge fusion in otolaryngology monitoring can be realized.
[0097] In summary, the technical solution adopted in this application can improve the image quality of endoscopic monitoring images, and generate the correlation map of patients in combination with the monitoring text to realize the knowledge fusion in otolaryngology monitoring.
[0098] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems) and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowcharts and / or block diagrams can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowcharts and / or block diagrams can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0099] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.
[0100] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
Claims
1. An otolaryngology intelligent diagnosis and monitoring system based on multimodal data fusion, characterized in that, The described diagnosis and monitoring system includes: An image acquisition module, which is used to acquire multi-modal monitoring data of a patient during otolaryngology monitoring, and extract the endoscopic monitoring images of the patient from the multi-modal monitoring data; An image restoration module, which is used to perform edge monitoring on the endoscopic monitoring images to obtain endoscopic monitoring edge images, generate a secretion perturbation mask based on the endoscopic monitoring edge images, and determine the endoscopic monitoring restored images of the patient according to the secretion perturbation mask and the endoscopic monitoring images; An information decision-making module, which is used to determine the pixel decision windows of each pixel point in the endoscopic monitoring restored images, perform information decision-making on each pixel point in the endoscopic monitoring restored images based on the corresponding pixel decision windows, and then obtain the endoscopic monitoring decision images of the patient. Specifically, it includes: Determine the pathological background noise of the endoscopic monitoring restored images, where the pathological background noise represents the background noise level in the endoscopic monitoring restored images, calculate the background noise variance of the endoscopic monitoring restored images, and use the calculation result as the pathological background noise of the endoscopic monitoring restored images; For each pixel point in the endoscopic monitoring restored images, determine the local information complexity and local pixel level of the pixel point according to the corresponding pixel decision window. Among them, the local information complexity represents the complexity of the pixel information contained in the pixel decision window corresponding to the pixel point, and the variance of all pixel points contained in the pixel decision window corresponding to the pixel point is used as the local information complexity of the pixel point. The local pixel level represents the average pixel value contained in the pixel decision window corresponding to the pixel point; Perform information decision-making on the pixel point through the pathological background noise, the local information complexity, and the local pixel level to obtain the pixel value of the decision pixel point corresponding to the pixel point, and then obtain the decision pixel points corresponding to each pixel point in the endoscopic monitoring restored images; Among them, the pixel value of the decision pixel point corresponding to the pixel point is determined by the following formula: , represents the pixel value of the decision pixel corresponding to the pixel point , represents the pixel value of the pixel point . represents the local information complexity of the pixel point . represents the local pixel level of the pixel point . represents the pathological background noise; Combine all decision pixel points according to the corresponding image coordinates to obtain the endoscopic monitoring decision images of the patient; An associated atlas generation module, which is used to extract the monitoring text data of the patient from the multi-modal monitoring data, perform feature fusion on the monitoring text data and the endoscopic monitoring decision images, and then generate a fusion associated atlas of the patient during otolaryngology monitoring.
2. The intelligent diagnosis and monitoring system for otorhinolaryngology based on multi-modal data fusion according to claim 1, wherein The multi-modal monitoring data includes endoscopic monitoring images and monitoring text data of the patient during otolaryngology monitoring.
3. An otolaryngology intelligent diagnosis and monitoring system based on multi-modal data fusion as claimed in claim 1, characterized in that, Use the Canny edge detection algorithm to perform edge detection on the endoscopic monitoring images.
4. An otolaryngology intelligent diagnosis and monitoring system based on multimodal data fusion according to claim 1, characterized in that, Specifically, generating a secretion perturbation mask based on the endoscopic monitoring edge images includes: Perform morphological processing on the endoscopic monitoring edge images to obtain closed edge images; Perform hole filling on the closed edge images to obtain single-connected region images; Perform erosion processing on the single-connected region images, and generate a secretion perturbation mask based on the erosion processing results.
5. An intelligent diagnosis and monitoring system for otolaryngology based on multimodal data fusion according to claim 1, characterized in that, Determining the endoscopic monitoring restoration image of the patient based on the secretion perturbation mask and the endoscopic monitoring image is to perform an exclusive OR operation on the secretion perturbation mask and the endoscopic monitoring image, and then obtain the endoscopic monitoring restoration image of the patient.
6. The intelligent diagnosis and monitoring system for otorhinolaryngology based on multimodal data fusion according to claim 1, characterized in that, Determining the pixel decision window of each pixel point in the endoscopic monitoring restoration image specifically includes: Obtaining a preset pixel window; For each pixel point in the endoscopic monitoring restoration image, determining the information fineness corresponding to the pixel point according to the pixel window; Based on the information fineness, determining the pixel decision window of the pixel point, and then obtaining the pixel decision window of each pixel point in the endoscopic monitoring restoration image.
7. An otolaryngology intelligent diagnosis and monitoring system based on multi-modal data fusion according to claim 6, characterized in that, Determining the information fineness corresponding to the pixel point according to the pixel window specifically includes: Taking the pixel point as the window center of the pixel window, and then determining all neighborhood windows of the pixel window; Determining the pixel heterogeneity coefficient between the pixel window and each neighborhood window; Determining the information fineness corresponding to the pixel point through all the pixel heterogeneity coefficients.
8. An intelligent diagnosis and monitoring system for otolaryngology based on multimodal data fusion according to claim 6, characterized in that, Determining the pixel decision window of the pixel point based on the information fineness specifically includes: Obtaining a preset fineness threshold; When the information fineness is greater than the fineness threshold, setting the pixel decision window size of the pixel point to 3 pixels × 3 pixels; When the information fineness is not greater than the fineness threshold, setting the pixel decision window size of the pixel point to 7 pixels × 7 pixels.
9. An intelligent diagnosis and monitoring system for otolaryngology based on multi-modal data fusion as claimed in claim 1, wherein Performing feature fusion on the monitoring text data and the endoscopic monitoring decision image, and then generating a fusion correlation map of the patient in otolaryngology monitoring specifically includes: Performing semantic parsing on the monitoring text data to obtain the structured knowledge of the patient in otolaryngology monitoring; Performing lesion feature extraction on the endoscopic monitoring decision image to obtain a lesion image feature vector; Performing joint modeling on the structured knowledge and the lesion image feature vector, and then generating a fusion correlation map of the patient in otolaryngology monitoring.
Citation Information
Patent Citations
Image processing method and device of endoscope image, terminal and readable storage medium
CN115965603A
Laryngoscope image multi-attribute classification method based on multi-modal information fusion
CN116664929A