Newborn ear malformation intelligent diagnosis method and system based on multi-modal fusion
Through the multimodal fusion method, combined with the improved YOLOv1 and SwinTransformer network, the geometric and texture features of the auricle image are extracted, and the CBAM and adaptive feature fusion modules are used to solve the problem of insufficient accuracy of the single-stage detection model in the diagnosis of neonatal ear deformities, and achieve higher diagnostic accuracy.
Patent Information
- Application Number
- CN202510692744.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-05
AI Technical Summary
The single-stage detection YOLO model in existing technologies has limited ability to capture multi-scale features, cannot fully capture local details of the ear, and has difficulty distinguishing morphologically similar neonatal ear deformities, resulting in low diagnostic accuracy.
A multimodal fusion method was adopted, combined with an improved YOLOv1 detection network and a SwinTransformer classification network. The geometric and texture features of the auricle images were extracted, and feature fusion was performed using the CBAM module and the adaptive feature fusion module to construct a diagnostic model for neonatal ear deformities.
It improves the ability to capture local details of the ear, enhances the recognition of complex morphology and texture differences, and improves the accuracy and robustness of the diagnosis of neonatal ear deformities.
Smart Images

Figure CN120600281A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and in particular to an intelligent diagnosis method and system for neonatal ear deformity based on multimodal fusion. Background Art
[0002] The ear is the primary organ for hearing, and ear deformities can lead to hearing loss, hearing impairment, or morphological and aesthetic defects. The neonatal period is a critical period for hearing and language development, as well as for non-invasive auricular correction. Early diagnosis of ear deformities facilitates timely medical intervention, preventing hearing and morphological defects from impacting children's language learning and social confidence.
[0003] Traditionally, the diagnosis of ear deformities relies primarily on physician experience, making uncommon deformities (such as missing helix crus and abnormally protruding concha) prone to misdiagnosis or missed diagnosis. With the rapid development of science and technology, and the urgent need for standardized diagnostic criteria for neonatal ear deformities, more and more modern technologies are being applied to the diagnosis of neonatal ear deformities.
[0004] Currently, image recognition technology using the single-stage detection YOLO model has been applied to the diagnosis of neonatal ear deformities. However, the single-stage detection YOLO model has limited ability to capture multi-scale features, cannot fully capture local details of the ear, and has difficulty distinguishing morphologically similar deformity subtypes, resulting in low accuracy in neonatal ear deformity diagnosis. Summary of the Invention
[0005] In view of the above shortcomings of the existing technology, the purpose of the embodiments of the present invention is to provide an intelligent diagnosis method for neonatal ear deformities based on multimodal fusion, which can solve the technical problems existing in the existing technology that the single-stage detection YOLO model has limited ability to capture multi-scale features, cannot fully capture local details of the ear, has difficulty in distinguishing morphologically similar deformity subtypes, and has low accuracy in diagnosing neonatal ear deformities.
[0006] A first aspect of an embodiment of the present invention provides an intelligent diagnosis method for neonatal ear deformity based on multimodal fusion, comprising:
[0007] S1: Acquire the auricle image of the newborn;
[0008] S2: Construct a diagnostic model for neonatal ear deformities;
[0009] S3: extracting geometric features of the auricle image and outputting an ear region detection frame in the auricle image;
[0010] S4: extracting texture features of the auricle image;
[0011] S5: performing multimodal fusion on the geometric features and the texture features to obtain fusion features of the auricle image;
[0012] S6: Detect whether the newborn has ear deformity based on the fusion features of the auricle image, and output the type of deformity.
[0013] A second aspect of an embodiment of the present invention provides an intelligent diagnosis system for neonatal ear deformity based on multimodal fusion, comprising: a processor and a memory;
[0014] The memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the intelligent diagnosis method for neonatal ear deformity based on multimodal fusion as described in the first aspect are implemented.
[0015] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0016] In an embodiment of the present invention, the geometric features and texture features in the newborn auricle image can be multimodally fused to fully capture the local details of the ear, better capture the complex morphological and texture differences, improve the model's ability to distinguish different types of ear deformities, and improve the accuracy of neonatal ear deformity diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings are only for the purpose of illustrating specific embodiments and are not to be considered as limiting the present invention. Throughout the drawings, the same reference symbols represent the same components. Obviously, the drawings described below are only some embodiments of the present invention. It is clear that those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0018] Figure 1 1 is a flow chart of an intelligent diagnosis method for neonatal ear deformity based on multimodal fusion provided by an embodiment of the present invention;
[0019] Figure 2 1 is a schematic structural diagram of an intelligent diagnosis method for neonatal ear deformity based on multimodal fusion provided by an embodiment of the present invention;
[0020] Figure 3 This is a schematic diagram of a cloud-edge collaboration architecture provided by an embodiment of the present invention;
[0021] Figure 4 This is a structural diagram of an intelligent diagnosis system for neonatal ear deformity based on multimodal fusion provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0022] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are part of the embodiments of the present invention, rather than all of the embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work should fall within the scope of protection of the present invention.
[0023] The following describes in detail the intelligent diagnosis method for neonatal ear deformity based on multimodal fusion provided by the embodiment of the present invention through specific embodiments and application scenarios in conjunction with the accompanying drawings.
[0024] Reference Manual Figure 1 , which shows a flow chart of an intelligent diagnosis method for neonatal ear deformity based on multimodal fusion provided by an embodiment of the present invention.
[0025] Reference Manual Figure 2 , shows a structural schematic diagram of an intelligent diagnosis method for neonatal ear deformity based on multimodal fusion provided by an embodiment of the present invention.
[0026] The present invention provides an intelligent diagnosis method for neonatal ear deformity based on multimodal fusion, which may include the following steps:
[0027] S1: Acquire an image of the newborn's pinna.
[0028] Specifically, the image of the newborn's auricle can be taken with a mobile phone, a professional camera, or related computer-aided equipment. The present invention does not limit the method of taking the auricle image.
[0029] S2: Construct a diagnostic model for neonatal ear deformities.
[0030] Specifically, the neonatal ear deformity diagnosis model may include: a YOLOv11 detection network, a SwinTransformer classification network, a CBAM module, and / or an adaptive feature fusion module. The specific contents of the neonatal ear deformity diagnosis model will be introduced in detail later.
[0031] S3: Extracting geometric features of the auricle image and outputting an ear region detection frame in the auricle image.
[0032] In a possible implementation, S3 specifically includes: extracting geometric features of the auricle image through an improved YOLOv11 detection network, and outputting an ear region detection frame in the auricle image.
[0033] Optionally, the YOLOv11 detection network includes: a backbone network, a neck network, and a detection head, and the YOLOv11 detection network is improved by replacing the original network structure of the backbone network with a lightweight GhostNet network structure.
[0034] The YOLOv11 detection network is an advanced single-stage object detection model designed to quickly and accurately identify target areas in images. It improves detection speed and accuracy through an improved network architecture and optimized algorithms.
[0035] It is important to note that the improved YOLOv11 detection network extracts geometric features from auricle images and outputs an ear region detection bounding box, effectively improving the accuracy and efficiency of ear deformity diagnosis. As an efficient object detection model, YOLOv11 can quickly locate the ear region and accurately extract relevant geometric features, providing high-quality input for subsequent deformity analysis. This approach not only improves detection speed, making it suitable for real-time applications, but also accurately identifies the auricle region in complex backgrounds, providing reliable support for subsequent diagnosis and treatment decisions.
[0036] In an embodiment of the present invention, by replacing the backbone network of the YOLOv11 detection network with a lightweight GhostNet network structure, the efficiency and processing speed of the model can be significantly improved while reducing the consumption of computing resources. GhostNet reduces the computational complexity and memory usage of the model while maintaining high accuracy by utilizing a lighter network structure, making the detection of the ear area faster and more efficient. Furthermore, it helps in deployment on the user terminal side (such as a mobile phone) when adopting a cloud-edge collaborative architecture.
[0037] S4: Extracting texture features of the auricle image.
[0038] In a possible implementation, the S4 specifically includes: extracting texture features of the auricle image through an improved Swin Transformer classification network.
[0039] Among them, the Swin Transformer classification network is an image classification model based on the Transformer architecture, which is designed to handle visual tasks through efficient local and global feature extraction.
[0040] Optionally, the Swin Transformer classification network includes multiple Swin Transformer blocks, and a local perception block is inserted before each Swin Transformer block to improve the Swin Transformer classification network.
[0041] It should be noted that the improved Swin Transformer classification network works specifically as follows:
[0042] The auricle image is input.
[0043] The auricle image is divided into non-overlapping image blocks by a patch partitioning layer.
[0044] The original pixel RGB values of each image block are concatenated into an image block vector.
[0045] Each of the image block vectors is reshaped into a spatial feature map through the local perception block.
[0046] The spatial feature map is convolved by 3×3 dilated convolution and GELU activation function, and the extraction capability of spatial local features is increased by residual connection. The spatial feature map is transformed into an intermediate feature map and input into the SwinTransformer block.
[0047] It is important to note that the local perception block image vector is reshaped into a spatial feature map and convolved using a 3×3 dilated convolution and a GELU activation function, effectively enhancing the extraction of local features. The dilated convolution expands the receptive field to capture a wider range of contextual information, while the GELU activation function helps improve the expressiveness of nonlinear features. Furthermore, the introduction of residual connections effectively avoids the vanishing gradient problem and improves the transfer efficiency of deep features, further enhancing the extraction of spatial local features.
[0048] The texture features of the auricle image are extracted through the window multi-head self-attention, shifted window multi-head self-attention and multi-layer perceptron in the Swin Transformer block:
[0049]
[0050] Among them, X l-1 represents the output of the l-1 layer, LN represents layer normalization, WMSA represents the window multi-head self-attention mechanism, represents the output of the window multi-head self-attention mechanism after adding the residual, MLP represents the multi-layer perceptron, X l represents the output of the lth layer, SWMSA represents the shift window multi-head self-attention mechanism, represents the output of the shift window multi-head self-attention mechanism after adding the residual, X l+1 Represents the output of the l+1th layer, and the output of the last layer is the extracted texture features.
[0051] It should be noted that, combined with the window multi-head self-attention mechanism in the Swin Transformer block, global and local texture features can be efficiently extracted, providing richer and more accurate texture information for subsequent multimodal fusion.
[0052] In an embodiment of the present invention, the texture features of auricle images are extracted using an improved Swin Transformer classification network, significantly improving the ability to recognize ear details and subtle texture changes. The Swin Transformer effectively captures both local and global information in an image through its local perception blocks and windowed self-attention mechanism, excelling particularly when processing complex textures. This improvement enables more accurate extraction of subtle texture features of the auricle, providing a richer and more precise basis for deformity diagnosis, thereby improving diagnostic accuracy and robustness.
[0053] S5: Perform multimodal fusion on the geometric features and the texture features to obtain fusion features of the auricle image.
[0054] In a possible implementation, the S5 specifically includes: performing multimodal fusion on the geometric features and the texture features through a CBAM module to obtain fusion features of the auricle image.
[0055] The Convolutional Block Attention Module (CBAM) is an attention mechanism designed to enhance the feature representation capabilities of convolutional neural networks. CBAM introduces channel-wise and spatial-wise attention mechanisms to weight feature maps in the channel and spatial dimensions, respectively, highlighting important features and suppressing irrelevant ones.
[0056] In this embodiment of the present invention, the introduction of the CBAM module for multimodal fusion of geometric and texture features effectively enhances the expressive power of the fused features. CBAM adaptively adjusts feature weights across channels and space, helping the model automatically focus on the most recognizable key information, thereby improving the accuracy and robustness of deformity feature extraction in auricle images. This approach enhances the network's ability to recognize complex deformity types, thus supporting more accurate diagnosis.
[0057] Furthermore, although the CBAM module helps the model automatically focus on the most recognizable key information by adaptively adjusting the feature weights in channels and space, it tends to fixedly enhance certain channels or spatial regions, and the network cannot flexibly select the most important information for fusion based on the characteristics of the specific input image. Therefore, the present invention proposes a new adaptive feature fusion method.
[0058] In a possible implementation, the S5 specifically includes:
[0059] The geometric features and the texture features are multimodally fused through an adaptive feature fusion module to obtain fusion features of the auricle image.
[0060] The specific fusion method of the adaptive feature fusion module is:
[0061] Downsampling is performed on a geometric feature map composed of the geometric features and a texture feature map composed of the texture features so that the sizes of the geometric feature map and the texture feature map are consistent.
[0062] Using the Softmax activation function, we can get the intermediate scores of the downsampled geometric feature map and texture feature map:
[0063]
[0064] in, Represents the median score of the geometric eigenvalue at position (u, v), W g represents the fusion weight matrix of the geometric feature map, Represents the geometric eigenvalue at position (u, v) in the geometric feature map, b g represents the fusion bias of the geometric feature map, Represents the median score of the texture feature value at position (u, v), W t represents the fusion weight matrix of the texture feature map, Represents the geometric eigenvalue at position (u, v) in the texture feature map, b t Represents the fusion bias term of the texture feature map, and Softmax represents the Softmax activation function.
[0065] It's important to note that intermediate scores play a crucial role in adaptive feature fusion, representing the "importance" or "attention" of the feature value at each location in the geometric and texture feature maps relative to other features. These intermediate scores determine which feature type (geometric or texture) contributes more to the final fusion result during the fusion process. By adaptively adjusting these scores, the model can flexibly select which features are more important in specific regions or contexts, leading to more precise feature fusion and final decision-making.
[0066] Calculate the adaptive weight coefficient according to the intermediate score between the geometric feature map and the texture feature map:
[0067]
[0068] in, Represents the adaptive weight coefficient of the geometric eigenvalue at position (u,v), Represents the adaptive weight coefficient of the texture feature value at position (u,v).
[0069] Perform multimodal fusion on the geometric feature map and the texture feature map:
[0070]
[0071] Among them, y uv Represents the fused feature value at position (u,v).
[0072] In an embodiment of the present invention, an adaptive feature fusion module fuses geometric and texture features, adaptively adjusting their weights based on the importance of the geometric and texture features at each location. This approach calculates the median score for each feature using a Softmax activation function and uses these scores to dynamically adjust the fusion ratio of geometric and texture features, ensuring that the most discernible features receive greater attention. This adaptive fusion approach can enhance the model's ability to recognize different features, thereby providing more accurate auricle image analysis results and improving the accuracy of ear deformity detection.
[0073] S6: Detect whether the newborn has ear deformity based on the fusion features of the auricle image, and output the type of deformity.
[0074] Among them, the types of deformities may include: drooping ears, monkey ears, helix deformity, retracted ears, hidden ears, protruding ears, cup-shaped ears, abnormal protrusion of the concha, etc.
[0075] Specifically, the fused features can be probability mapped through the Softmax activation function to determine the probability of various ear deformities in newborns, and then detect whether the newborns have ear deformities and output the type of deformity.
[0076] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:
[0077] In an embodiment of the present invention, the geometric features and texture features in the newborn auricle image can be multimodally fused to fully capture the local details of the ear, better capture the complex morphological and texture differences, improve the model's ability to distinguish different types of ear deformities, and improve the accuracy of neonatal ear deformity diagnosis.
[0078] Reference Manual Figure 3 , showing a schematic diagram of the cloud-edge collaboration architecture provided by an embodiment of the present invention.
[0079] In a possible embodiment, in a possible embodiment, the S2 is specifically: using a cloud-edge collaborative architecture to build a newborn ear deformity diagnosis model. A lightweight YOLOv11 detection network is deployed on the user terminal side to extract the geometric features of the auricle image and output the ear area detection frame in the auricle image. A Swin Transformer classification network, a CBAM module and / or an adaptive feature fusion module are deployed on the edge side to extract the texture features of the auricle image, perform multimodal fusion on the geometric features and the texture features, obtain the fusion features of the auricle image, and detect whether the newborn has ear deformities based on the fusion features of the auricle image. Dynamic federated learning is coordinated in the cloud, and the newborn ear deformity diagnosis model stored on the user terminal and the edge side is trained through dynamic federated learning.
[0080] In an embodiment of the present invention, a cloud-edge collaborative architecture is used to construct a neonatal ear deformity diagnosis model, which can efficiently combine the computing resources of user terminals, edge devices, and the cloud. By deploying a lightweight YOLOv11 detection network on the user terminal, the geometric features of the auricle image can be quickly extracted and the ear area can be located, reducing dependence on remote servers and improving real-time performance and response speed. Deploying the Swin Transformer classification network, CBAM module, and adaptive feature fusion module on the edge side can accurately extract and fuse texture features, enhancing the ability to recognize ear deformities.
[0081] Furthermore, hospital data has high requirements for privacy. Dynamic federated learning and model training in the cloud not only ensures data privacy and reduces the risk of data leakage, but also continuously optimizes the performance of diagnostic models.
[0082] Optionally, the user terminal side is specifically a user's mobile phone, the edge side is specifically a private server of the hospital, and the cloud side is specifically a cloud server.
[0083] In one possible implementation, the specific method of training the neonatal ear deformity diagnosis model using dynamic federated learning based on homomorphic encryption includes:
[0084] On each of the edge sides, the auricle image dataset stored on the edge side is input into the neonatal ear deformity diagnosis model for detection, the detection result is compared with the true label (the true label and detection frame of each auricle image training sample is determined by professionals for diagnosis), and the loss function value is calculated (this application uses the mean square error loss function, and the cross entropy loss function, etc. can also be used), and then the gradient of the model stored on the edge side is calculated according to the loss function value:
[0085]
[0086] Among them, G ik represents the gradient of the i-th edge side, J i Represents the mean square error loss function value on the i-th edge side, θ k represents the network parameters at the kth iteration.
[0087] The gradient on the edge side is homomorphically encrypted and uploaded to the cloud.
[0088] The cloud decrypts the gradient on the edge side.
[0089] Aggregate the decrypted gradients:
[0090]
[0091] Among them, G k represents the aggregate gradient at the kth iteration, λ i Represents the weight coefficient of the i-th edge side.
[0092] Update the global model parameters using the gradient descent method:
[0093] θ k+1 =θ k ―η k G k +β·(θ k ―θ k―1 )
[0094] Among them, θ k+1 represents the model parameters at the k+1th iteration, η k represents the adaptive learning rate at the kth iteration, and β represents the momentum coefficient.
[0095] It's important to note that updating global model parameters through gradient descent and introducing a momentum term can accelerate model convergence and improve training stability. By combining the current gradient with the gradient from the previous update, the momentum coefficient helps the model escape from local optima and accelerate its movement toward the global optimal solution, thereby improving training efficiency, reducing training time, and enhancing final model performance. This approach is particularly effective with large-scale data and complex tasks.
[0096] The adaptive learning rate is calculated as:
[0097]
[0098] Among them, η k represents the adaptive learning rate at the kth iteration, η0 represents the initial learning rate, G u represents the aggregate gradient at the u-th iteration, u=1,2,…,k, and μ represents a hyperparameter to prevent the denominator from being zero.
[0099] It's important to note that the adaptive learning rate calculation method adjusts the learning rate at each iteration based on the cumulative sum of squared gradients. This means that during training, when the gradient is large, the learning rate automatically decreases to avoid oscillation or instability caused by excessive step sizes. When the gradient is small, the learning rate increases appropriately to promote faster convergence. This method effectively and dynamically adjusts the learning rate, improving training stability and efficiency.
[0100] The cloud sends global model parameters to each edge side, so that each edge side updates its stored neonatal ear deformity diagnosis model according to the global model parameters.
[0101] In an embodiment of the present invention, the dynamic federated learning training model for neonatal ear deformity diagnosis based on homomorphic encryption can fully utilize distributed computing resources while ensuring data privacy. The gradient is calculated on each edge side and uploaded to the cloud for homomorphic encryption to ensure that private information is not leaked during data transmission. The cloud decrypts and aggregates the encrypted gradients, updates the global model parameters using the gradient descent method, and then sends the updated global model parameters to each edge side for model update. This process effectively realizes collaborative learning between multiple devices. Through the calculation and dynamic adjustment of the adaptive learning rate, the model training process can converge faster and more stably, avoiding the problems of overfitting and slow convergence. This method not only improves the training efficiency of the model, but also ensures data privacy protection, so that efficient and secure learning and model optimization can be carried out in a distributed environment.
[0102] In a possible implementation, performing homomorphic encryption on the edge-side gradient specifically includes:
[0103] Set the ciphertext modulus q, key distribution χ, error distribution ψ, and generate the random vector α for this round of transmission.
[0104] Each edge side generates a private key based on the key distribution, generates an error vector based on the error distribution, and calculates the public key of each edge side:
[0105] b i =-s i a+e i (modq)
[0106] Among them, b i represents the public key of the ith edge side, s i represents the private key of the ith edge, α represents a random vector, e i represents the error vector on the i-th edge side, mod represents the modulo operation, and q represents the ciphertext modulus.
[0107] All devices collaborate to calculate the aggregated public key:
[0108]
[0109] in, Represents an aggregate public key.
[0110] Using the aggregated public key, the gradient of the local model is encrypted as plaintext to obtain the ciphertext:
[0111]
[0112] Among them, ct i represents the ciphertext on the i-th edge side, c 0i represents the first ciphertext component on the i-th edge side, c 1i represents the second ciphertext component on the i-th edge side, v i represents the random polynomial on the ith edge side sampled from the key distribution, m i represents the plaintext on the ith edge side, that is, the gradient of the model on the ith edge side, and β represents the public parameter.
[0113] Send the ciphertext to the server.
[0114] In an embodiment of the present invention, by homomorphically encrypting the gradients on the edge side, secure model training can be performed without exposing the local data of the device. By generating a private key, a public key and collaboratively calculating the aggregated public key, each edge side encrypts the gradient of its model and sends it to the server, thereby ensuring data privacy and security. During the encryption process, the gradient is transmitted in ciphertext form to prevent leakage or tampering in the middle, while also ensuring that the model can still retain valid information when performing global aggregation. The advantage of doing this is that it not only protects data privacy and prevents the leakage of sensitive information, but also enables each edge device to safely participate in joint learning, avoiding the security risks brought about by centralized storage or transmission of data.
[0115] In a possible implementation, the cloud decrypting the edge-side gradient specifically includes:
[0116] The server calculates the aggregate ciphertext:
[0117]
[0118] Among them, C sum Represents aggregate ciphertext, C sum0 Represents the first aggregate ciphertext component, C sum1 Represents the second aggregate ciphertext component.
[0119] The server broadcasts the second aggregate ciphertext component to all devices.
[0120] Each edge side calculates its decryption share:
[0121]
[0122] Among them, D i represents the decryption share of the i-th edge side, represents the independent noise term generated by the i-th edge signal, which is used to further enhance privacy. This error term is tiny and can be eliminated by rounding off after subsequent homomorphic operations.
[0123] Each edge side sends the decrypted share to the server.
[0124] The server combines all decrypted shares with the first aggregate ciphertext component and recovers the plaintext:
[0125]
[0126] Among them, M represents the aggregated plaintext, which contains the gradients of the models on each edge side.
[0127] In an embodiment of the present invention, through this decryption and aggregation process based on homomorphic encryption, the server can securely calculate and recover the gradients uploaded by each edge device while ensuring that the privacy of the device's local data is not leaked. During the decryption process, each device calculates its own decryption share and sends it to the server. The server merges these decryption shares with the rest of the aggregated ciphertext to recover the gradient of the model. This method effectively avoids the risk of directly exposing the device's local data, while further enhancing data privacy protection by introducing tiny independent noise terms. Through this encryption and decryption process, it can be ensured that data privacy is fully protected during dynamic federated learning, while achieving secure cross-device collaboration and global model updates.
[0128] The intelligent diagnosis method for neonatal ear deformities based on multimodal fusion provided in the embodiments of the present application can be executed by an intelligent diagnosis device for neonatal ear deformities based on multimodal fusion. In the embodiments of the present application, the intelligent diagnosis method for neonatal ear deformities based on multimodal fusion is performed by an intelligent diagnosis device for neonatal ear deformities based on multimodal fusion as an example to illustrate the intelligent diagnosis device for neonatal ear deformities based on multimodal fusion provided in the embodiments of the present application.
[0129] Reference Manual Figure 4 , shows a structural schematic diagram of a neonatal ear deformity intelligent diagnosis system based on multimodal fusion provided by an embodiment of the present invention.
[0130] The embodiment of the present invention provides a newborn ear deformity intelligent diagnosis system 20 based on multimodal fusion, comprising: a processor 201 and a memory 202;
[0131] The memory 202 stores a program or instruction that can be run on the processor 201. When the program or instruction is executed by the processor 201, the steps of the above-mentioned intelligent diagnosis method for neonatal ear deformity based on multimodal fusion are implemented, and the same technical effect can be achieved. To avoid repetition, the present invention will not be repeated.
[0132] It should be understood that the processor 201 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0133] It should also be understood that the memory 202 in the embodiment of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0134] The above embodiments can be implemented in whole or in part through software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0135] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0136] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0137] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0138] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another edge side, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0139] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0140] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0141] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0142] An embodiment of the present invention provides a readable storage medium comprising: a program or instruction stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the above-mentioned intelligent diagnosis method for neonatal ear deformity based on multimodal fusion are implemented, and the same technical effect can be achieved. To avoid repetition, the present invention will not be repeated.
[0143] Finally, it should be noted that the above embodiments are merely illustrative of the technical solutions of the embodiments of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they may still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein with equivalents; and such modifications or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention. Any changes or replacements that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be covered by the scope of protection of the present invention.
Claims
1. An intelligent diagnosis method for neonatal ear deformity based on multimodal fusion, characterized in that: include: S1: Acquire the auricle image of the newborn; S2: Construct a diagnostic model for neonatal ear deformities; S3: extracting geometric features of the auricle image and outputting an ear region detection frame in the auricle image; S4: extracting texture features of the auricle image; S5: performing multimodal fusion on the geometric features and the texture features to obtain fusion features of the auricle image; S6: Detect whether the newborn has ear deformity based on the fusion features of the auricle image, and output the type of deformity.
2. The intelligent diagnosis method for neonatal ear deformity based on multimodal fusion according to claim 1, characterized in that: The S3 is specifically: Extracting geometric features of the auricle image through an improved YOLOv11 detection network, and outputting an ear region detection frame in the auricle image; The YOLOv11 detection network includes: backbone network, neck network, and detection head. The YOLOv11 detection network is improved by replacing the original network structure of the backbone network with the lightweight GhostNet network structure.
3. The intelligent diagnosis method for neonatal ear deformity based on multimodal fusion according to claim 2, characterized in that: The S4 specifically comprises: extracting texture features of the auricle image through an improved Swin Transformer classification network; The Swin Transformer classification network includes multiple Swin Transformer blocks, and a local perception block is inserted before each Swin Transformer block to improve the Swin Transformer classification network.
4. The intelligent diagnosis method for neonatal ear deformity based on multimodal fusion according to claim 3, characterized in that: The S5 specifically includes: The geometric features and the texture features are multimodally fused through the CBAM module to obtain fused features of the auricle image.
5. The intelligent diagnosis method for neonatal ear deformity based on multimodal fusion according to claim 3, characterized in that: The S5 specifically includes: Performing multimodal fusion on the geometric features and the texture features through an adaptive feature fusion module to obtain fusion features of the auricle image; The specific fusion method of the adaptive feature fusion module is: Downsampling a geometric feature map composed of the geometric features and a texture feature map composed of the texture features so that the sizes of the geometric feature map and the texture feature map are consistent; Use the Softmax activation function to obtain the intermediate scores of the downsampled geometric feature map and texture feature map; Calculating an adaptive weight coefficient according to an intermediate score between the geometric feature map and the texture feature map; Multimodal fusion is performed on the geometric feature map and the texture feature map.
6. The intelligent diagnosis method for neonatal ear deformity based on multimodal fusion according to claim 5, characterized in that: The S2 is specifically: Using a cloud-edge collaborative architecture, a diagnostic model for neonatal ear deformities was built; Deploying a lightweight YOLOv11 detection network on the user terminal side to extract geometric features of the auricle image and output an ear region detection frame in the auricle image; Deploying a Swin Transformer classification network, a CBAM module, and / or an adaptive feature fusion module on the edge to extract texture features of the auricle image, performing multimodal fusion on the geometric features and the texture features to obtain fused features of the auricle image, and detecting whether the newborn has ear deformities based on the fused features of the auricle image; Dynamic federated learning is coordinated in the cloud to train the neonatal ear deformity diagnosis model stored in the user terminal and the edge side through dynamic federated learning.
7. The intelligent diagnosis method for neonatal ear deformity based on multimodal fusion according to claim 6, characterized in that: The user terminal side is specifically the user's mobile phone, the edge side is specifically the hospital's private server, and the cloud side is specifically the cloud server.
8. The intelligent diagnosis method for neonatal ear deformity based on multimodal fusion according to claim 6, characterized in that: The specific method of training the neonatal ear deformity diagnosis model using dynamic federated learning based on homomorphic encryption includes: At each edge side, using the auricle image dataset stored at the edge side, calculating the gradient of the model stored at the edge side; Homomorphically encrypt the gradient on the edge side and upload it to the cloud; The cloud decrypts the gradient on the edge side; Aggregate the decrypted gradients; Update global model parameters through gradient descent method; The cloud sends global model parameters to each edge side, so that each edge side updates its stored neonatal ear deformity diagnosis model according to the global model parameters.
9. The intelligent diagnosis method for neonatal ear deformity based on multimodal fusion according to claim 8, characterized in that: The homomorphic encryption of the edge-side gradient specifically includes: Each edge side generates a private key based on the key distribution, generates an error vector based on the error distribution, and calculates the public key of each edge side; All devices collaboratively calculate the aggregated public key; Using the aggregated public key, encrypt the gradient of the local model as plaintext to obtain ciphertext; Send the ciphertext to the server; The cloud-side decryption of the edge-side gradient specifically includes: The server calculates the aggregate ciphertext; The server broadcasts the second aggregate ciphertext component to all devices; Each edge side calculates its decryption share; Each edge side sends the decrypted share to the server; The server combines all decrypted shares with the first aggregate ciphertext component and recovers the plaintext.
10. An intelligent diagnosis system for neonatal ear deformity based on multimodal fusion, characterized in that: include: processor and memory; The memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the intelligent diagnosis method for neonatal ear deformity based on multimodal fusion as described in claims 1 to 9 are implemented.
Citation Information
Cited By
Newborn ear reconstruction and malformation diagnosis method and system based on improved three-dimensional Gaussian splashing
CN121616785A
Ear disease diagnosis method based on adaptive attention fusion and boundary optimization network
CN121885163A