Artificial cochlea intelligent debugging method based on multi-modal artificial intelligence
By integrating objective detection data and subjective feedback through multimodal artificial intelligence methods, the parameters of the cochlear implant are automatically optimized, solving the problem of insufficient parameter adjustment accuracy in existing technologies and achieving efficient and convenient parameter adjustment and improved hearing experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-05
- Publication Date
- 2026-06-09
AI Technical Summary
Current cochlear implant parameter adjustments rely on in-person operations by professional audiologists, which cannot match the patient's dynamic electro-auditory needs in real time and fails to effectively integrate the patient's subjective auditory feedback, resulting in insufficient parameter fitting accuracy, affecting the wearing experience and auditory rehabilitation effects.
Employing a multimodal artificial intelligence approach, this method acquires multimodal information from users, including objective detection data, the user's environment, and subjective feedback. It then uses a multimodal neural network model to fuse and process this information, automatically outputting and optimizing cochlear implant parameters, and supporting iterative optimization until the user is satisfied.
It achieves the fusion analysis of multimodal information, improves the accuracy and convenience of parameter adjustment, supports patients to participate in the adjustment, adapts to multiple scenarios, improves the auditory experience, and can determine the cause of changes in auditory perception, thereby improving the effectiveness of parameter adjustment.
Smart Images

Figure CN122163994A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an intelligent adjustment method for cochlear implants based on multimodal artificial intelligence. Background Technology
[0002] Currently, cochlear implant parameter adjustment relies on in-person operations by professional audiologists, primarily depending on their subjective assessment. This requires manually setting core parameters such as T-value (threshold), C-value (comfort value), pulse width, and stimulation rate, based on results from pure-tone audiometry, hearing aid assessment, and speech recognition threshold tests. This approach suffers from drawbacks such as a high dependence on experienced professionals for adjustment cycles and an inability to match patients' dynamic electroacoustic needs in real-time. Furthermore, existing adjustment systems often rely solely on objective test data, failing to integrate the patient's subjective auditory feedback, resulting in insufficient parameter fitting accuracy and impacting the wearing experience and auditory rehabilitation outcomes. Even after fitting, it is difficult for users to determine the cause of changes in their daily auditory perception, hindering effective adjustments to parameters such as noise reduction level and directionality. Summary of the Invention
[0003] This invention provides an intelligent cochlear implant tuning method based on multimodal artificial intelligence to solve at least one of the above problems.
[0004] In a first aspect, embodiments of the present invention provide an intelligent cochlear implant tuning method based on multimodal artificial intelligence, comprising: S110. Obtain the user's multimodal information, which includes objective detection data, the current scene, existing cochlear implant parameters, and the user's existing subjective feedback on the existing cochlear implant parameters. S120. The multimodal information is fused using a multimodal neural network model to automatically output new cochlear implant parameters; S130. If the user's new subjective feedback on the new cochlear implant parameters is unsatisfactory, construct new multimodal information based on the new cochlear implant parameters and the new subjective feedback, and return to S120 based on the new multimodal information until the user obtains cochlear implant parameters that satisfy the user.
[0005] In a second aspect, embodiments of the present invention provide an electronic device, the electronic device comprising: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the intelligent cochlear implant debugging method based on multimodal artificial intelligence as described in any embodiment.
[0006] Thirdly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the intelligent cochlear implant tuning method based on multimodal artificial intelligence as described in any embodiment.
[0007] In summary, this embodiment provides a multimodal artificial intelligence-based intelligent cochlear implant adjustment method. It integrates multimodal detection data acquisition, objective auditory feedback recording, intelligent parameter calculation, and autonomous adjustment interaction. It can achieve multi-dimensional fusion analysis based on eABR, eSRT, eCAP, CAEP, MMN test results and patient auditory perception, automatically generating and optimizing all-dimensional parameters of the cochlear implant. Simultaneously, it supports the participation of novice audiologists and patients in parameter adjustments, improving fitting accuracy and ease of use. Specifically, this embodiment can achieve the following beneficial effects: 1. By integrating multimodal objective test data with patients' subjective auditory feedback, the accuracy of parameter tuning is improved compared to traditional manual tuning, making it more suitable for individual patient needs; 2. It supports patients to independently complete parameter adjustments without relying on professional audiologists for in-person operation, shortening the adjustment cycle and improving ease of use; 3. Adapt to various environmental parameters and automatically match noise reduction and directional modes in different scenarios to improve the listening experience in complex environments; 4. After the user obtains satisfactory cochlear implant parameters, any changes in auditory perception during daily life can help determine whether the changes are due to normal scene changes or changes in the user's hearing function. Different parameter adjustment strategies can then be adopted for different causes, improving the effectiveness of parameter adjustment while protecting the patient's hearing. Attached Figure Description
[0008] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0009] Figure 1 This is a flowchart of an intelligent cochlear implant tuning method based on multimodal artificial intelligence provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a multimodal neural network model provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0011] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0012] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0013] Figure 1 This is a flowchart illustrating an intelligent cochlear implant adjustment method based on multimodal artificial intelligence, provided by an embodiment of the present invention. This method is used to intelligently adjust the cochlear implant parameters of a cochlear implant wearer and is executed by an electronic device. Figure 1 As shown, the method specifically includes: S110. Obtain the user's multimodal information, which includes objective detection data, the current scene, existing cochlear implant parameters, and the user's existing subjective feedback on the existing cochlear implant parameters.
[0014] Objective test data can be collected at a cochlear implant fitting center using specialized equipment. Optionally, the user's raw objective test data can be collected first, including at least one of the following: eABR (electro-Auditory Brainstem Evoked Potentials), eSRT (electro-Stokes Reflex Threshold), eCAP (electro-Evoked Auditory Neuromuscular Compound Action Potential), CAEP (Auditory Cortical Evoked Potentials), and MMN (Auditory Mismatch Negative Wave Test). This embodiment supports structured input of raw test data and direct import of data from the device.
[0015] In one specific implementation, the data format of eSTR can include two types: 1. Time-domain recording: x-axis represents time (ms), and y-axis represents acoustic impedance value (mmho). 2. Intensity-response curve: x-axis represents electrical stimulation intensity (mA), and y-axis represents change in acoustic impedance (mmho).
[0016] eCAP data can be presented in two formats: 1. Waveform graph: x-axis represents time (ms), y-axis represents voltage (μV). 2. Intensity-amplitude graph: x-axis represents stimulus current intensity (mA), y-axis represents eCAP amplitude (μV).
[0017] CAEP data can be presented in two formats: 1. Waveform graph: x-axis represents time (ms), y-axis represents voltage (μV). 2. Intensity-amplitude graph: x-axis represents sound intensity (dB), y-axis represents P1 amplitude (μV).
[0018] MMN data can be presented as a waveform graph, with the x-axis representing time (ms) and the y-axis representing voltage (μV).
[0019] eABR data can be presented in two formats: 1. Waveform graph: x-axis represents time (ms), y-axis represents voltage (μV). 2. Intensity-amplitude graph: x-axis represents stimulus current intensity (mA), y-axis represents wave V amplitude (μV).
[0020] Then, cochlear implant fitting-related test indicators can be extracted from the original objective test data as the final objective test data. Optionally, the above original test data can be normalized and averaged using a sliding window, and standardized features such as the wave V latency of eABR, the electrical stimulation speech recognition threshold of eSRT, the latency and amplitude of CAEP, and the amplitude of MMN can be extracted to form an objective test feature library.
[0021] Optionally, the cochlear implant parameters include at least one of the following: Basic electrical stimulation parameters: T-value (threshold), C-value (comfort value), stimulation pulse width, stimulation rate, volume reference value, sensitivity level; Acoustic processing parameters: outdoor unit noise reduction level, microphone signal noise reduction processing intensity, microphone directionality mode, and automatic environmental recognition threshold.
[0022] Optionally, the user's scenario can include types such as home, workplace, and noisy environment, and subjective feedback can include auditory descriptions of volume, sound clarity, and background noise interference. In one specific implementation, a visual interactive interface can be designed for the user, providing standardized auditory perception description options (such as volume too loud / too soft, sound blurry / clear, background noise interference, etc.), supporting patients to customize their auditory experience in specific scenarios (home, workplace, noisy environment), and converting subjective feedback into quantitative feature data.
[0023] S120. The multimodal information is fused using a multimodal neural network model to automatically output new cochlear implant parameters.
[0024] This step uses deep learning algorithms to construct a multimodal neural network model, which is used to fuse multimodal information and automatically generate cochlear implant parameters.
[0025] In one specific implementation, the following can be used: Figure 2 The network model shown includes a feature encoding layer, a cross-modal feature fusion layer, and a parameter prediction layer. The entire data processing flow includes the following steps: Step 1: Embedded encoding of each modality's information is performed using a feature encoding layer. Optionally, the feature encoding layer can use an MLP (Multilayer Perceptron) to process the final objective detection data. And existing cochlear implant parameters Convert to 3D eigenvectors and The scene is defined by embedding. and subjective feedback characteristics Convert to 3D eigenvectors and ,in For natural numbers greater than 1:
[0026]
[0027]
[0028]
[0029] Step 2: Through a cross-modal feature fusion layer, the encoded features of each modality are fused to obtain multimodal features. Optionally, objective test data (objective indicators) reflect the user's hearing function status and are the main basis for determining cochlear implant parameters. The user's environment and subjective feedback are also important auxiliary factors in hearing aid fitting, both of which are related to objective test data and thus affect the cochlear implant wearing experience. To consider this correlation simultaneously, this embodiment adopts the following... Figure 2 The network structure shown constructs bidirectional attention weights for "objective indicators-environment" and "objective indicators-subjective feedback" through dual-path cross-attention layers to facilitate better multimodal information fusion. For ease of distinction and description, this embodiment refers to these two cross-attention layers as the first cross-attention layer and the second cross-attention layer, respectively.
[0030] Among them, the query feature Query in the first cross-attention layer is the scene ( Both the Key and Value are feature vectors composed of objective detection data and existing cochlear implant parameters (denoted as...). More specifically, The dimension is , The dimension is Then you can get The attention matrix is used to further calculate the fused multimodal features. The entire operation of the cross-attention layer can be represented as follows:
[0031] Similarly, the query feature Query in the second cross-attention layer is subjective feedback ( The key and value are still a feature vector composed of objective detection data and existing cochlear implant parameters. More specifically, The dimension is , The dimension is Then you can get The attention matrix is used to further calculate the fused multimodal features. The entire operation of the cross-attention layer can be represented as follows:
[0032] For ease of distinction and description, this embodiment will... and These are referred to as the first multimodal feature and the second multimodal feature, respectively.
[0033] Step 3: Process the multimodal features using the parameter prediction layer to output new cochlear implant parameters. This step automatically calculates and outputs the full-dimensional parameters of the cochlear implant based on the multimodal fusion results described above. Optionally, the parameter prediction layer can use MLP_HEAD, taking the parameters obtained in the previous step... and Feature vectors are transformed into a full-dimensional parameter set for cochlear implants. :
[0034] S130. If the user's new subjective feedback on the new cochlear implant parameters is unsatisfactory, construct new multimodal information based on the new cochlear implant parameters and the new subjective feedback, and return to S120 based on the new multimodal information until the user obtains cochlear implant parameters that satisfy the user.
[0035] This embodiment supports iterative optimization of cochlear implant parameters, dynamically adjusting parameter values based on continuous user feedback on auditory perception, gradually approaching the optimal fit. Optionally, after generating cochlear implant parameters, if the patient is not satisfied with the generated parameters, a new set of cochlear implant parameters can be generated based on the patient's subjective feedback and the currently generated parameters. This process is repeated until a satisfactory hearing experience is achieved for the user.
[0036] Furthermore, the method of this embodiment can be integrated into a terminal such as an APP or a cochlear implant external controller, or it can be deployed on a backend server or other electronic devices. Once a user obtains satisfactory cochlear implant parameters at a cochlear implant fitting center, they can wear it in daily life. If the user's daily environment and / or subjective hearing experience (i.e., subjective feedback) change, the user can input the changed information through the APP or other terminal. The method in the terminal, backend server, or other electronic device will then regenerate the appropriate cochlear implant parameters based on the changed information and the remaining unchanged information, and apply them to the user's cochlear implant.
[0037] However, since users cannot update objective test data in their daily lives, the objective test data used in user-initiated adjustments are all data from the previous cochlear implant fitting center. If the user's hearing function changes during wear, leading to poor subjective feedback, the objective test data may not match the user's actual hearing function, resulting in the inability to generate suitable cochlear implant parameters. Therefore, this embodiment provides a method: after obtaining satisfactory cochlear implant parameters, if the parameters need to be regenerated due to changes in the user's environment and / or subjective feedback, the method uses the attention distribution of the changed information and the unchanged objective test data to determine whether the objective test data needs updating, thereby effectively adjusting the cochlear implant parameters. In one specific implementation, this process may include the following steps: Step 1: After obtaining satisfactory cochlear implant parameters for the user, when the user's scenario and / or subjective feedback change, new multimodal information is constructed based on the changed scenario and / or subjective feedback, as well as other unchanged information. The other unchanged information includes at least objective test data and existing cochlear implant parameters. If only one of the scenario or subjective feedback changes, the other unchanged information also includes the unchanged aspect of either the scenario or the subjective feedback.
[0038] Step 2: Using the first cross-attention layer, an attention distribution map is calculated using the scene in the new multimodal information as the query feature and the objective detection data as the key and value features. Simultaneously, using the second cross-attention layer, another attention distribution map is calculated using the scene's subjective feedback in the new multimodal information as the query feature and the objective detection data as the key and value features. Combined with... Figure 2 In this embodiment, the data processing branch shown by the dashed line in the figure is enabled, and two attention distribution maps are output. For ease of distinction and description, they are referred to as the first attention distribution map and the second attention distribution map, respectively. Optionally, in conjunction with the foregoing description, the first cross-attention layer can obtain... The attention matrix, from By extracting the dimension corresponding to the part of the objective detection data, we can obtain... The first attention distribution map represents the correlation between the surrounding environment and the objective detection data. Similarly, the second cross-attention layer... By extracting the portion corresponding to the objective detection data from the attention matrix, we can obtain... The second attention distribution map represents the correlation between subjective feedback and objective detection data.
[0039] Step 3: Based on the first and second attention distribution maps, determine whether the objective detection data is at risk of failure. Since the user has already obtained satisfactory cochlear implant parameters, if the user's subjective feedback changes, this change may stem from a change in the scene or a change in the user's hearing function. If the user's scene changes, then assuming the user's hearing function remains unchanged, this scene change should bring about a certain change in the second subjective feedback. In summary: after the user has obtained satisfactory cochlear implant parameters, under normal circumstances, scene changes and changes in subjective feedback should follow and match each other. The corresponding objective detection data should be able to explain both changes simultaneously, or in other words, the channels that the scene focuses on in the objective detection data and the channels that the subjective feedback should focus on in the objective detection data should match each other. For example, if high-frequency noise increases in a scene, the objective test data should focus on 8kHz ASSR (high frequency). Therefore, the user's feedback should also focus on 8kHz ASSR (high frequency) in the objective test data. This is consistent with the influence of the environment on the user's feedback. If the user's feedback of "muffled" sounds like "dull" but the objective test data focuses on 2kHz ASSR (low frequency), it shows a contradiction and mismatch. This may be because the user's hearing function has changed, and the current objective test data can no longer reflect the user's actual hearing condition.
[0040] Optionally, the two attention distribution maps mentioned above can be used to characterize the interpretation of objective detection data by the scene and subjective feedback. In one specific implementation, the scene information can first be treated as a whole, and the first attention distribution map can be aggregated to obtain the attention distribution curves of the overall scene and each dimension of the objective detection data. That is, for The first attention distribution map is summed and averaged across all rows to obtain... The attention distribution curve is obtained. Simultaneously, the subjective feedback can be treated as a whole, and the second attention distribution map can be aggregated to obtain the attention distribution curves of the overall subjective feedback and each dimension of the objective detection data. For ease of distinction and description, this embodiment refers to the two curves as the first attention distribution curve and the second attention distribution curve, respectively.
[0041] Then, the first and second attention distribution curves are normalized respectively. The similarity between the two normalized curves is determined based on the Euclidean distance and the intersection-union ratio (IU / R) of hotspot regions. Optionally, continuous curve segments exceeding a certain multiple of the average value can be identified as hotspot regions. The IU / R of hotspot regions in the two curves is then calculated. This IU / R reflects the dimensional difference between the scene and the objective detection data focused on by subjective feedback, thus reflecting the similarity of the curves. Simultaneously, Euclidean distance is also a measure of curve similarity; a weighted average of the two similarities yields a comprehensive similarity.
[0042] If the overall similarity is greater than a set threshold, it indicates that the scene and subjective feedback have a high degree of similarity with the objective detection data, and it can be determined that the possibility of the objective detection data being invalid is extremely small, with almost no risk. However, if the overall similarity is less than or equal to the set threshold, it indicates that the relationship between the scene and subjective feedback and the objective detection data is very different, which does not conform to the natural law that subjective feedback and scene changes should follow and match each other when the user's hearing function remains unchanged. In this case, the objective detection data is at risk of being invalid and may no longer reflect the user's current actual hearing function. For the sake of simplicity, in this embodiment, the above two situations will be referred to as no risk of invalidation and risk of invalidation, respectively.
[0043] Step 4: If there is no risk of failure, fuse the multimodal information and output new cochlear implant parameters based on the fusion features; if there is a risk of failure, update the objective detection data. Specifically, if the objective detection data has no risk of failure, the subsequent operations of the cross-modal feature fusion layer and parameter prediction layer can continue to be performed to obtain new cochlear implant parameters. If the objective detection data has a risk of failure, it needs to be updated. For example, the user may be prompted to go to a cochlear implant fitting center to undergo relevant hearing tests again so that the new objective detection data can be used to regenerate the cochlear implant parameters. In practical applications, if the objective detection data has a risk of failure, it may be impossible to obtain good cochlear implant parameters no matter how it is adjusted, or even if a barely acceptable cochlear implant parameter is obtained, it may still have potential damage to the patient's hearing. Therefore, it is still necessary to prompt the user to undergo hearing tests again.
[0044] Accordingly, to achieve the above objectives, during model training, expert-annotated cochlear implant parameters can be used as supervisory labels. Supervised training methods are then employed to continuously approximate the labeled cochlear implant parameters in the overall model output. Next, separate training sets and loss functions are constructed for the first and second attention distribution maps, and the model is fine-tuned using the loss function from supervised training.
[0045] In one specific implementation, after obtaining cochlear implant parameters that satisfy the user: If the user's hearing function remains unchanged, subjective feedback changes due to scene changes are collected, and a positive sample is formed by the objective detection data under the hearing function, the scene after each change, and the subjective feedback; if the user's hearing function changes, subjective feedback changes due to scene changes are collected, and a negative sample is formed by the objective detection data before the change, the scene after each change, and the subjective feedback. This yields training sets for the first and second attention distribution maps.
[0046] Then, the first and second cross-attention layers are fine-tuned using the positive and negative samples to ensure that the similarity differences between the first and second attention maps corresponding to the positive and negative samples can be distinguished. Optionally, the following loss function can be constructed:
[0047] in, The similarity is represented by the number of positive samples. The similarity is represented by the negative samples. This represents the similarity threshold between 0 and 1; express The identifier function, hour , hour , This represents the sum of the identification functions of all positive samples in the same batch; express The identifier function, hour , hour , This represents the sum of the label functions of all negative samples in the same batch. (The rest of the text appears to be a fragment and requires further context for accurate translation.) Maximizing this allows the similarity between positive and negative samples to be passed through a threshold. Correspondingly, this threshold can also serve as a basis for subsequent failure risk assessment.
[0048] The model is fine-tuned by using this loss function (or other loss functions with the same function) together with the original loss function in the previous supervision method (i.e. the model output and the labeled information have the smallest difference).
[0049] It should be noted that the positive and negative samples in the above fine-tuning were constructed using data after obtaining cochlear implant parameters that satisfied the user. Figure 2 The section marked with a dotted line only applies to situations where the user has obtained satisfactory cochlear implant parameters but then experiences changes in the environment and / or subjective feedback. It is unnecessary for parameter adjustments performed immediately after objective indicator testing at the fitting center. This is because the user's hearing function generally does not change during the short period at the fitting center, and the existing cochlear implant parameters at this time are not yet stable enough to provide a stable hearing effect. The user's subjective feedback is not entirely due to changes in the environment; it may also be caused by discomfort with the existing cochlear implant parameters. Therefore, a large difference between the first and second attention distribution maps is normal and does not indicate that the objective test data is invalid. However, once the user returns to daily life after obtaining satisfactory parameters at the fitting center, if they experience changes in hearing and adjust their subjective feedback, and / or change the environment, then this step can be performed.Figure 2 The portion shown by the dashed line can be used with the aid of two attention distribution maps to determine whether objective test data may have changed. In this implementation, it is possible to determine whether changes in a user's auditory perception in daily life stem from normal scene changes or from changes in the user's hearing function, and to employ different parameter adjustment strategies for different causes, thereby improving the effectiveness of parameter adjustment while protecting the patient's hearing.
[0050] In summary, this embodiment provides a multimodal artificial intelligence-based intelligent cochlear implant adjustment method. It integrates multimodal detection data acquisition, objective auditory feedback recording, intelligent parameter calculation, and autonomous adjustment interaction. It can achieve multi-dimensional fusion analysis based on eABR, eSRT, eCAP, CAEP, MMN test results and patient auditory perception, automatically generating and optimizing all-dimensional parameters of the cochlear implant. Simultaneously, it supports the participation of novice audiologists and patients in parameter adjustments, improving fitting accuracy and ease of use. Specifically, this embodiment can achieve the following beneficial effects: 1. By integrating multimodal objective test data with patients' subjective auditory feedback, the accuracy of parameter tuning is improved compared to traditional manual tuning, making it more suitable for individual patient needs; 2. It supports patients to independently complete parameter adjustments without relying on professional audiologists for in-person operation, shortening the adjustment cycle and improving ease of use; 3. Adapt to various environmental parameters and automatically match noise reduction and directional modes in different scenarios to improve the listening experience in complex environments; 4. After a user obtains satisfactory cochlear implant parameters, the system can help determine whether objective test data is at risk of failure when the user's environment and / or subjective feedback change. Specifically, when used to assess changes in auditory perception in daily life, it can help determine whether changes are more likely due to normal scene changes or changes in the user's hearing function, allowing for different parameter adjustment strategies to be adopted for different causes, improving the effectiveness of parameter adjustment while protecting the patient's hearing.
[0051] It should be noted that all user data involved in this application is information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0052] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, such as... Figure 3 As shown, the device includes a processor 60, a memory 61, an input device 62, and an output device 63; the number of processors 60 in the device can be one or more. Figure 3Taking a processor 60 as an example; the processor 60, memory 61, input device 62, and output device 63 in the device can be connected via a bus or other means. Figure 3 Taking the example of a connection between China and Israel via a bus.
[0053] The memory 61, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the intelligent cochlear implant tuning method based on multimodal artificial intelligence in this embodiment of the invention. The processor 60 executes various functional applications and data processing of the device by running the software programs, instructions, and modules stored in the memory 61, thereby realizing the aforementioned intelligent cochlear implant tuning method based on multimodal artificial intelligence.
[0054] The memory 61 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 61 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory, or other non-volatile solid-state storage device. In some instances, the memory 61 may further include memory remotely located relative to the processor 60, which can be connected to the device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0055] Input device 62 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the device. Output device 63 may include display devices such as a display screen.
[0056] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements any embodiment of the intelligent cochlear implant tuning method based on multimodal artificial intelligence.
[0057] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0058] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0059] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0060] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as C or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0061] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A method for intelligent adjustment of cochlear implants based on multimodal artificial intelligence, characterized in that, include: S110. Obtain the user's multimodal information, which includes objective detection data, the current scene, existing cochlear implant parameters, and the user's existing subjective feedback on the existing cochlear implant parameters. S120. The multimodal information is fused using a multimodal neural network model to automatically output new cochlear implant parameters; S130. If the user's new subjective feedback on the new cochlear implant parameters is unsatisfactory, construct new multimodal information based on the new cochlear implant parameters and the new subjective feedback, and return to S120 based on the new multimodal information until the user obtains cochlear implant parameters that satisfy the user.
2. The method according to claim 1, characterized in that, S110 includes: Obtain the user's original objective test data, wherein the original objective test data includes at least one of the following: electrical stimulation auditory brainstem evoked potential test data, electrical stimulation stapedius muscle reflex threshold test data, electrical stimulation evoked auditory nerve compound action potential test data, auditory cortex evoked potential test data, and auditory mismatch negative wave test test data; From the original objective test data, relevant test indicators for cochlear implant fitting are extracted as the final objective test data.
3. The method according to claim 1, characterized in that, Cochlear implant parameters include: at least one basic electrical stimulation parameter, and / or, at least one acoustic processing parameter.
4. The method according to claim 1, characterized in that, S120 includes: Each modality information is embedded and encoded separately through a feature coding layer; By using a cross-modal feature fusion layer, the encoded features of information from each modality are fused to obtain multimodal features; The multimodal features are processed using a parameter prediction layer to output new cochlear implant parameters.
5. The method according to claim 1, characterized in that, After obtaining satisfactory cochlear implant parameters from the user, the following steps are also included: When the user's context and / or subjective feedback change, new multimodal information is constructed based on the changed context and / or subjective feedback, as well as other unchanged information. Through the first cross-attention layer, using the scene in the new multimodal information as the query feature and the objective detection data as the key feature and value feature, a first attention distribution map is calculated; Through the second cross-attention layer, the subjective feedback in the new multimodal information is used as the query feature, and the objective detection data is used as the key feature and value feature to calculate the second attention distribution map; Based on the first attention distribution map and the second attention distribution map, determine whether the objective detection data has a risk of failure; If there is no risk of failure, the multimodal information is fused, and new cochlear implant parameters are output based on the fusion features. If there is a risk of failure, update the objective test data.
6. The method according to claim 5, characterized in that, The step of determining whether the objective detection data has a risk of failure based on the first attention distribution map and the second attention distribution map includes: Aggregate the first attention distribution map to obtain the first attention distribution curves of the overall scene and each dimension of the objective detection data; Aggregate the second attention distribution map to obtain the overall subjective feedback and the second attention distribution curves of each dimension of the objective detection data; Normalize the first attention distribution curve and the second attention distribution curve respectively; The similarity between the two normalized curves is determined by the Euclidean distance and the intersection-union ratio of hotspot regions. If the similarity is greater than a set threshold, it is determined that the objective detection data has no risk of failure.
7. The method according to claim 5, characterized in that, Before calculating the first attention distribution map using the scene in the new multimodal information as the query feature and the objective detection data as the key and value features through the first cross-attention layer, the method further includes: After obtaining satisfactory cochlear implant parameters from the user: Without changing the user's hearing function, subjective feedback changes caused by scene changes are collected, and a positive sample is formed by the objective test data under the hearing function, the scene after each change, and the subjective feedback. When a user's hearing function changes, the subjective feedback changes due to the change in the scenario, and a negative sample is formed by the objective test data before the change, the scenario after each change, and the subjective feedback. The first and second cross-attention layers are trained using positive and negative samples, so that the similarity differences between the first and second attention maps corresponding to positive and negative samples can be distinguished.
8. The method according to claim 5, characterized in that, The method is applicable to situations where objective detection data cannot be updated in real time.
9. An electronic device, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the intelligent cochlear implant debugging method based on multimodal artificial intelligence as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the intelligent cochlear implant tuning method based on multimodal artificial intelligence as described in any one of claims 1-8.