Multi-mode monitoring method, device and equipment for glottis change and medium

By setting up an audio and video collector at the patient's oropharyngeal cavity, combining image and audio recognition models, the laryngeal mask leak is solved in real time, and the problem of difficulty in monitoring the laryngeal mask leak during anesthesia is improved, and the patient's safety is improved.

CN120147690APending Publication Date: 2025-06-13FANG ZHIHUA (BEIJING) MEDICAL TECHNOLOGY DEVELOPMENT CENTER (LLP)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510139659.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

During anesthesia, air leakage of the laryngeal mask is more common, which can easily lead to the patient's ventilation and the increased risk of asphyxiation, and may cause postoperative complications such as abdominal distension, nausea, and vomiting. It is crucial to monitor air leakage of the laryngeal mask in real time.

Method used

The audio and video collector set at the target position of the patient's oropharyngeal cavity can obtain audio data and image data, and use the pre-trained target image recognition model and target audio recognition model to analyze the data to determine whether there is a risk of air leakage in the laryngeal mask.

Benefits of technology

Real-time and accurate monitoring of air leakage in the laryngeal mask is achieved, dangerous situations caused by air leakage are avoided, and patient safety is improved during anesthesia.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147690A_ABST
    Figure CN120147690A_ABST
Patent Text Reader

Abstract

The invention provides a glottis change multi-mode monitoring method, device and equipment and a medium, and the method comprises the steps: obtaining audio data and image data of an oropharyngeal cavity of a patient through an audio and video collector disposed at a target position of the oropharyngeal cavity of the patient; performing data processing on the audio data to obtain target audio data with target features; respectively inputting the target audio data and the image data into a pre-trained target image recognition model and a pre-trained target audio recognition model, and respectively determining an audio recognition risk value and an image recognition risk value of the patient; according to the audio recognition risk value and the image recognition risk value, whether the laryngeal mask air leakage risk exists or not is determined. The effect of accurately monitoring air leakage of the laryngeal mask in real time is achieved, and danger caused by air leakage of the laryngeal mask in the anesthesia process of a patient is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of medical data monitoring. Specifically, it relates to a multimodal monitoring method, device, equipment, and medium for glottal changes. Background Technique

[0002] With the improvement of medical standards and the update of the concept of perioperative management, especially the establishment of the laryngeal mask airway is simpler, more convenient, has fewer complications, and the patient experience is better. Therefore, the laryngeal mask is more and more widely used in the perioperative period.

[0003] However, using a laryngeal mask also has certain risks. After the establishment of the laryngeal mask airway, it is very common for laryngeal mask leakage to occur during anesthesia maintenance. Once leakage occurs, it will affect the patient's ventilation. In more serious cases, the situation of inability to intake air may occur, leading to asphyxia risks. At the same time, laryngeal mask leakage is extremely likely to cause the airflow of positive pressure ventilation to enter the esophagus, resulting in gastric distension, which can affect the operation of intra-abdominal surgery and is also likely to cause postoperative abdominal distension, nausea, vomiting, and aspiration. Therefore, how to monitor laryngeal mask leakage during anesthesia maintenance is extremely important for patients undergoing general anesthesia with a laryngeal mask. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a multimodal monitoring method, device, equipment, and medium for glottal changes, which can obtain the audio data and image data of the patient's oropharynx through an audio-video collector set at the target position of the patient's oropharynx, and analyze the audio data and image data through a target image recognition model and a target audio recognition model to determine whether there is a risk of laryngeal mask leakage, achieving the effect of monitoring laryngeal mask leakage in real time and accurately, and avoiding risks caused by laryngeal mask leakage during the patient's anesthesia.

[0005] In the first aspect, an embodiment of this application provides a multimodal monitoring method for glottal changes. The method includes: obtaining the audio data and image data of the patient's oropharynx through an audio-video collector set at the target position of the patient's oropharynx; performing data processing on the audio data to obtain target audio data with target features; respectively inputting the target audio data and the image data into a pre-trained target image recognition model and a pre-trained target audio recognition model to respectively determine the audio recognition risk value and the image recognition risk value of the patient; and determining whether there is a risk of laryngeal mask leakage according to the audio recognition risk value and the image recognition risk value.

[0006] Optionally, the step of performing data processing on the audio data to obtain target audio data with target features includes: determining first intermediate audio data from the audio data according to a preset target sound intensity range; and determining target audio data from the first intermediate audio data according to a preset target audio range.

[0007] Optionally, the audio recognition model includes multiple voiceprint matching recognition models, and each voiceprint matching model is used to recognize a voiceprint feature; wherein, each voiceprint matching model is trained through the following steps: obtaining a plurality of training audio data corresponding to the voiceprint matching model; inputting each training audio data into the voiceprint matching model to obtain a training voiceprint matching value; calculating a loss function value of the voiceprint matching model according to the training voiceprint matching value and the actual voiceprint matching value; training the voiceprint matching model according to the loss function value of the voiceprint matching model until the loss function value of the voiceprint matching model converges to the minimum value; determining the voiceprint matching model when the loss function value converges to the minimum value as the trained voiceprint matching model.

[0008] Optionally, the step of inputting each training audio data into the voiceprint matching model to obtain a training voiceprint matching value includes: performing noise reduction processing on the training audio data to determine the voiceprint of the training audio data; separating noise from the voiceprint according to the waveform of the voiceprint of the training audio data to obtain a target voiceprint waveform; matching the target voiceprint waveform with a preset standard voiceprint waveform to obtain a training voiceprint matching value.

[0009] Optionally, the method further includes: determining an audio recognition risk value of the patient according to the characteristic parameters of each voiceprint feature preset and the voiceprint matching value calculated by each voiceprint matching model; wherein, the characteristic parameters of each voiceprint feature correspond to the occurrence frequency of each voiceprint feature when there is a leak in the laryngeal mask.

[0010] Optionally, the method further includes: obtaining a plurality of body data of the patient; determining a correction parameter corresponding to the patient according to the plurality of body data; determining a plurality of standard voiceprint waveforms of the patient according to the correction parameter and a plurality of standard voiceprint waveforms; inputting the plurality of standard voiceprint waveforms of the patient into a pre-trained target audio recognition model, so that the target audio recognition model determines an audio recognition risk value according to the plurality of standard voiceprint waveforms of the patient.

[0011] Optionally, the method further includes: when installing the laryngeal mask, obtaining respiratory audio data and image data of the patient's oropharynx; respectively inputting the respiratory audio data and image data of the patient's oropharynx into a pre-trained target image recognition model and a pre-trained target audio recognition model to determine whether there is an abnormality in the installation of the laryngeal mask.

[0012] In a second aspect, an embodiment of the present application further provides a multimodal monitoring device for glottal changes, and the device includes:

[0013] A data acquisition module, configured to acquire audio data and image data of the patient's oropharynx through an audio-video collector disposed at a target position of the patient's oropharynx;

[0014] A data processing module for processing the audio data to obtain target audio data with target features;

[0015] A risk value calculation module for respectively inputting the target audio data and the image data into a pre-trained target image recognition model and a pre-trained target audio recognition model to respectively determine an audio recognition risk value and an image recognition risk value of the patient;

[0016] A leakage risk determination module for determining whether there is a laryngeal mask air leakage risk according to the audio recognition risk value and the image recognition risk value.

[0017] In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the above-mentioned multi-modal monitoring method for glottal changes are executed.

[0018] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the above-mentioned multi-modal monitoring method for glottal changes are executed.

[0019] The multi-modal monitoring method, device, equipment, and medium for glottal changes provided by the embodiments of the present application obtain audio data and image data of a patient's oropharynx through an audio-video collector disposed at a target position in the patient's oropharynx, and analyze the audio data and the image data through a target image recognition model and a target audio recognition model to determine whether there is a laryngeal mask air leakage risk, achieving the effect of monitoring the laryngeal mask air leakage in real time and accurately, and avoiding risks caused by laryngeal mask air leakage during the anesthesia of the patient.

[0020] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specific preferred embodiments are given in conjunction with the accompanying drawings and described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a flowchart of a multi-modal monitoring method for glottal changes provided by an embodiment of the present application;

[0023] Figure 2 Schematic diagram of the installation position of the audio - video collector provided by the embodiment of the present application;

[0024] Figure 3 Schematic diagram of the display for the audio recognition risk value and the image recognition risk value provided by the embodiment of the present application;

[0025] Figure 4 Schematic diagram of the structure of a multi - modal monitoring device for glottal changes provided by the embodiment of the present application;

[0026] Figure 5 Schematic diagram of the structure of an electronic device provided by the embodiment of the present application. Detailed implementation manners

[0027] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, every other embodiment obtained by those skilled in the art without creative efforts belongs to the scope of protection of the present application.

[0028] First, the applicable application scenarios of the present application are introduced. The present application can be applied to the technical field of medical data monitoring.

[0029] Through research, it is found that with the improvement of medical standards and the update of the concept of peri - operative management, especially the establishment of the laryngeal mask airway is simpler, more convenient, has fewer complications, and the patient experience is better. Therefore, the laryngeal mask is more and more widely used in the peri - operative period.

[0030] However, using a laryngeal mask also has certain risks. After the laryngeal mask airway is established, it is very common for laryngeal mask leakage to occur during anesthesia maintenance. Once leakage occurs, it will affect the patient's ventilation. In a more serious case, the situation of air intake failure may occur, leading to asphyxia risks. At the same time, laryngeal mask leakage is extremely likely to cause the airflow of positive - pressure ventilation to enter the esophagus, resulting in gastric distension, which can affect the operation of intra - abdominal surgery and is also likely to cause postoperative abdominal distension, nausea, vomiting, and aspiration. Therefore, how to monitor laryngeal mask leakage during anesthesia maintenance is extremely important for patients undergoing general anesthesia with a laryngeal mask.

[0031] As an example, during anesthesia maintenance, the position of the vocal cords varies with the degree of muscle relaxation, resulting in different sizes of glottis openings. The recovery of muscle relaxation causes the glottis to close, leading to an increase in inspiratory resistance. When the airway pressure is greater than the leakage pressure of the laryngeal mask, the laryngeal mask will leak, thus triggering the risk of asphyxia and various complications.

[0032] Based on this, the embodiments of the present application provide a multimodal monitoring method, device, equipment and medium for glottis changes, which can obtain the audio data and image data of the patient's oropharynx through an audio-video collector arranged at the target position of the patient's oropharynx, and analyze the audio data and image data through a target image recognition model and a target audio recognition model to determine whether there is a leakage risk of the laryngeal mask, achieving the effect of monitoring the leakage of the laryngeal mask in real time and accurately, and avoiding the danger caused by the leakage of the laryngeal mask during the anesthesia of the patient.

[0033] Please refer to Figure 1 , Figure 1 which is a flowchart of a multimodal monitoring method for glottis changes provided by an embodiment of the present application. As shown in Figure 1 , the multimodal monitoring method for glottis changes provided by the embodiments of the present application includes:

[0034] S101. Obtain the audio data and image data of the patient's oropharynx through an audio-video collector arranged at the target position of the patient's oropharynx.

[0035] As an example, please refer to Figure 2 , Figure 2 which is a schematic diagram of the installation position of the audio-video collector provided by an embodiment of the present application. As shown in Figure 2 , the schematic diagram of the installation position of the audio-video collector provided by the embodiments of the present application includes: a laryngeal mask 201, an air delivery tube 202, an audio-video collector installation position 303, a data transmission line 304, and a monitoring display 305.

[0036] In this way, the audio-video collector arranged at the oropharynx of the patient wearing the laryngeal mask can collect video data and audio data more clearly, and can display the situation in the patient's oropharynx in real time through the monitoring display.

[0037] S102. Perform data processing on the audio data to obtain target audio data with target features.

[0038] Specifically, the steps of performing data processing on the audio data to obtain target audio data with target features include: determining first intermediate audio data from the audio data according to a preset target sound intensity range; and determining target audio data from the first intermediate audio data according to a preset target audio range.

[0039] Here, it should be noted that in the case of laryngeal mask leakage, the patient will exhibit characteristics such as glottis closure, air bubbles passing through the periphery of the laryngeal mask body, and increased secretions in the oropharynx.

[0040] Among them, the pre-set target sound intensity range is the possible sound intensity range of the characteristics such as glottis closure, air bubbles passing through the periphery of the laryngeal mask body, and increased secretions in the oropharynx that occur in the patient determined through experiments.

[0041] The pre-set target audio range is the possible audio range of the characteristics such as glottis closure, air bubbles passing through the periphery of the laryngeal mask body, and increased secretions in the oropharynx that occur in the patient determined through experiments.

[0042] In this way, through the processing of audio data, it is possible to avoid transmitting redundant audio data to the audio recognition model, reduce the computational amount of the audio recognition model, and improve the recognition speed.

[0043] S103. Respectively input the target audio data and the image data into a pre-trained target image recognition model and a pre-trained target audio recognition model, and respectively determine the audio recognition risk value and the image recognition risk value of the patient.

[0044] Specifically, the audio recognition model includes multiple voiceprint matching recognition models, and each voiceprint matching model is used to recognize a kind of voiceprint feature.

[0045] Among them, each voiceprint matching model is trained through the following steps: obtaining a plurality of training audio data corresponding to the voiceprint matching model; inputting each training audio data into the voiceprint matching model to obtain a training voiceprint matching value; calculating the loss function value of the voiceprint matching model according to the training voiceprint matching value and the actual voiceprint matching value; training the voiceprint matching model according to the loss function value of the voiceprint matching model until the loss function value of the voiceprint matching model converges to the minimum value; determining the voiceprint matching model when the loss function value converges to the minimum value as the trained voiceprint matching model.

[0046] Specifically, the step of inputting each training audio data into the voiceprint matching model to obtain a training voiceprint matching value includes: performing noise reduction processing on the training audio data to determine the voiceprint of the training audio data; separating the noise from the voiceprint according to the waveform of the voiceprint of the training audio data to obtain a target voiceprint waveform; matching the target voiceprint waveform with a pre-set standard voiceprint waveform to obtain a training voiceprint matching value.

[0047] As an example, the multiple voiceprint matching recognition models may include: a glottis closure voiceprint matching recognition model, a voiceprint matching recognition model for the rupture of air bubbles appearing around the laryngeal mask body, a voiceprint matching recognition model for increased secretions in the oropharynx, and so on.

[0048] In this way, the present application can identify and match various possible voiceprints, and then determine the voiceprint matching value.

[0049] As an example, abnormal airflow sounds or gargling sounds caused by secretions can be monitored. Through audio signal processing, specific frequency and waveform features are identified by a voiceprint matching recognition model for voiceprint matching. The closure of the glottis or airway stenosis can be monitored. Through audio signal processing, specific frequency and waveform features are identified by a voiceprint matching recognition model for voiceprint matching. The bursting sound of bubbles can be monitored. Through audio signal processing, specific frequency and waveform features are identified by a voiceprint matching recognition model for voiceprint matching. When the matching degree of the voiceprint reaches a certain value within a period of time, it can be determined that laryngeal mask leakage is likely to occur. Combining with the video data, early warning of laryngeal mask leakage can be carried out.

[0050] Specifically, an open-source image semantic segmentation model (such as PP-LiteSeg) can be used for recognition. Image semantic segmentation models usually use convolutional neural networks (CNNs) to learn the complex relationships between pixels and perform segmentation. It uses an encoder-decoder structure. By learning a certain amount of annotated images, it can output the category of each pixel. For glottis detection, that is, identifying the pixel points in the glottis area of the input image.

[0051] For example, first, select some glottis pictures under a certain visual angle of the laryngeal mask for image semantic annotation to distinguish the edges of the vocal cords and the glottis. Then, use the annotated dataset as a training sample to train based on the image semantic segmentation model PP-LiteSeg. The generated model file is used for real-time video monitoring and calculation.

[0052] Exemplarily, the calculation logic of the image recognition model is as follows: Take the video frame after the anesthesiologist places the laryngeal mask and adjusts the position of the audio-video collector as the reference value of the normal glottis size. Identify the glottis area through the image semantic segmentation model, and count the number of pixel points in this target area as P0. The AI image semantic segmentation model continuously monitors and calculates each frame of the video image, calculates the number of pixel points Pt in the corresponding glottis area respectively, and calculates the ratio Rt of P0 / Pt at the same time. Rt is the image recognition risk value.

[0053] Specifically, the method further includes: determining the audio recognition risk value of the patient according to the characteristic parameters of each voiceprint feature set in advance and the voiceprint matching value calculated by each voiceprint matching model; wherein, the characteristic parameters of each voiceprint feature correspond to the occurrence frequency of each voiceprint feature when the laryngeal mask leaks.

[0054] Wherein, the method further includes: obtaining a plurality of physical data of a patient; determining a correction parameter corresponding to the patient according to the plurality of physical data; determining a plurality of standard voiceprint waveforms of the patient according to the correction parameter and a plurality of standard voiceprint waveforms; inputting the plurality of standard voiceprint waveforms of the patient into a pre-trained target audio recognition model, so that the target audio recognition model determines an audio recognition risk value according to the plurality of standard voiceprint waveforms of the patient.

[0055] Specifically, a multiple linear regression model can be used to calculate the correction parameter, and the variables include age, height, weight, mentohyoid distance, mouth opening, length of the hemimandible, etc.

[0056] As an example, each correction parameter can be calculated by the following formula:

[0057] S = α[β 1 A + β 2 B + …];

[0058] Wherein, s is a correction parameter matching a kind of voiceprint data, AB… are variables such as age, height, weight, mentohyoid distance, mouth opening, length of the hemimandible, etc., a calibration parameter of a kind of voiceprint is formed according to a plurality of data, and the calibration parameter is multiplied by the standard voiceprint waveform of this kind of voiceprint data to obtain a standard voiceprint waveform of the patient.

[0059] Specifically, the coefficient (β value) can be calculated through the following steps: First, sufficient patient sample data can be collected, including the actual situation of air leakage and the above-mentioned physiological parameters; clean the data and perform preprocessing to ensure that the data has no missing items and no outliers, etc.; secondly, use multiple linear regression or machine learning methods to perform regression analysis on the data to estimate the coefficients of each variable; finally, verify the accuracy of the model through the test set data to ensure the reliability and prediction ability of the model. After training, the coefficients (β values) of each physiological parameter can be obtained, which represent the influence degree of each parameter on the voiceprint. For example, if the coefficient of weight is relatively high, it means that weight has a greater influence on the generated voiceprint. Through these coefficients, the influence of different variables on the voiceprint size or fluctuation value can be quantitatively evaluated.

[0060] S104. Determine whether there is a risk of laryngeal mask air leakage according to the audio recognition risk value and the image recognition risk value.

[0061] As an example, please refer to Figure 3 , Figure 3Schematic diagram of the display for the audio recognition risk value and the image recognition risk value provided by the embodiments of the present application. The display content of the display for the audio recognition risk value and the image recognition risk value provided by the embodiments of the present application includes: the laryngeal mask monitoring video 305, the audio recognition risk value 303, and the image recognition risk value 304. The glottis 301 and the vocal cords 302 are displayed in the laryngeal mask monitoring video 305.

[0062] Optionally, the method further includes: when installing the laryngeal mask, acquiring the respiratory audio data and image data of the patient's oropharynx; respectively inputting the respiratory audio data and image data of the patient's oropharynx into a pre-trained target image recognition model and a pre-trained target audio recognition model to determine whether there is an abnormality in the installation of the laryngeal mask.

[0063] The multi-modal monitoring method for glottis changes provided by the embodiments of the present application can acquire the audio data and image data of the patient's oropharynx through an audio-video collector disposed at a target position in the patient's oropharynx, and analyze the audio data and image data through a target image recognition model and a target audio recognition model to determine whether there is a risk of air leakage in the laryngeal mask, achieving the effect of monitoring the air leakage of the laryngeal mask in real time and accurately, and avoiding the danger caused by the air leakage of the laryngeal mask during the anesthesia of the patient.

[0064] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a multi-modal monitoring of glottis changes provided by the embodiments of the present application. As shown in Figure 4 the multi-modal monitoring device 400 for glottis changes includes:

[0065] A data acquisition module 401, configured to acquire the audio data and image data of the patient's oropharynx through an audio-video collector disposed at a target position in the patient's oropharynx;

[0066] A data processing module 402, configured to perform data processing on the audio data to obtain target audio data with target features;

[0067] A risk value calculation module 403, configured to respectively input the target audio data and the image data into a pre-trained target image recognition model and a pre-trained target audio recognition model to respectively determine the audio recognition risk value and the image recognition risk value of the patient;

[0068] An air leakage risk determination module 404, configured to determine whether there is a risk of air leakage in the laryngeal mask according to the audio recognition risk value and the image recognition risk value.

[0069] The multi-modal monitoring device for glottis changes provided by the embodiments of the present application can obtain the audio data and image data of the patient's oropharynx through an audio-video collector disposed at the target position of the patient's oropharynx, and analyze the audio data and image data through a target image recognition model and a target audio recognition model to determine whether there is a risk of leakage of the laryngeal mask, achieving the effect of monitoring the leakage of the laryngeal mask in real time and accurately, and avoiding the danger caused by the leakage of the laryngeal mask during the anesthesia of the patient.

[0070] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided by the embodiments of the present application. As Figure 5 shown in

[0071] The memory 520 stores machine-readable instructions executable by the processor 510. When the electronic device 500 runs, the processor 510 communicates with the memory 520 through the bus 530. When the machine-readable instructions are executed by the processor 510, the steps of the multi-modal monitoring method for glottis changes in the method embodiment as described above can be executed. The specific implementation manner can refer to the method embodiment and will not be elaborated here. Figure 1 shown above.

[0072] The embodiments of the present application further provide a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of the multi-modal monitoring method for glottis changes in the method embodiment as described above can be executed. The specific implementation manner can refer to the method embodiment and will not be elaborated here. Figure 1 shown above.

[0073] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0074] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division manners in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the couplings, direct couplings, or communication connections shown or discussed with each other can be through some communication interfaces. The indirect couplings or communication connections between devices or units can be in electrical, mechanical, or other forms.

[0075] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0076] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist independently as individual physical units, or two or more units may be integrated in one unit.

[0077] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0078] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, and are not intended to limit them. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multimodal monitoring method for glottal changes, characterized in that: The method comprises: Acquiring audio data and image data of the patient's oropharyngeal cavity through an audio and video collector disposed at a target position of the patient's oropharyngeal cavity; Performing data processing on the audio data to obtain target audio data with target features; Inputting the target audio data and the image data into a pre-trained target image recognition model and a pre-trained target audio recognition model, respectively, to determine an audio recognition risk value and an image recognition risk value of the patient, respectively; Determine whether there is a risk of laryngeal mask air leakage based on the audio recognition risk value and the image recognition risk value.

2. The method according to claim 1, characterized in that The step of processing the audio data to obtain target audio data with target features includes: Determining first intermediate audio data from the audio data according to a preset target sound intensity range; According to a preset target audio range, target audio data is determined from the first intermediate audio data.

3. The method according to claim 1, characterized in that The audio recognition model includes a plurality of voiceprint matching recognition models, each voiceprint matching model is used to identify a voiceprint feature; Among them, each voiceprint matching model is trained through the following steps: Acquire multiple training audio data corresponding to the voiceprint matching model; Input each training audio data into the voiceprint matching model to obtain a training voiceprint matching value; Calculate the loss function value of the voiceprint matching model according to the training voiceprint matching value and the actual voiceprint matching value; Training the voiceprint matching model according to the loss function value of the voiceprint matching model until the loss function value of the voiceprint matching model converges to a minimum value; The voiceprint matching model when the loss function value converges to the minimum value is determined as the trained voiceprint matching model.

4. The method according to claim 3, characterized in that The steps of inputting each training audio data into the voiceprint matching model to obtain the training voiceprint matching value include: Performing noise reduction processing on the training audio data to determine the voiceprint of the training audio data; According to the waveform of the voiceprint of the training audio data, the voiceprint is subjected to noise separation to obtain a target voiceprint waveform; The target voiceprint waveform is matched with a preset standard voiceprint waveform to obtain a training voiceprint matching value.

5. The method according to claim 4, characterized in that The method further comprises: Determine the patient's audio recognition risk value based on the pre-set characteristic parameters of each voiceprint feature and the voiceprint matching value calculated by each voiceprint matching model; Among them, the characteristic parameter of each voiceprint feature corresponds to the frequency of occurrence of each voiceprint feature when the laryngeal mask air leaks.

6. The method according to claim 1, characterized in that The method further comprises: Acquire multiple physical data of the patient; determining a correction parameter corresponding to the patient based on the plurality of body data; Determining a plurality of standard voiceprint waveforms of the patient according to the calibration parameters and a plurality of standard voiceprint waveforms; The multiple standard voiceprint waveforms of the patient are input into a pre-trained target audio recognition model, so that the target audio recognition model determines the audio recognition risk value according to the multiple standard voiceprint waveforms of the patient.

7. The method according to claim 1, characterized in that The method further comprises: When the laryngeal mask is installed, the respiratory audio data and image data of the patient's oropharyngeal cavity are obtained; The breathing audio data and image data of the patient's oropharyngeal cavity are respectively input into a pre-trained target image recognition model and a pre-trained target audio recognition model to determine whether there is any abnormality in the installation of the laryngeal mask.

8. A multimodal monitoring device for glottal changes, characterized in that: The device comprises: A data acquisition module, used to acquire audio data and image data of the patient's oropharyngeal cavity through an audio and video collector arranged at a target position of the patient's oropharyngeal cavity; A data processing module, used for processing the audio data to obtain target audio data with target features; a risk value calculation module, used to input the target audio data and the image data into a pre-trained target image recognition model and a pre-trained target audio recognition model, respectively, to determine the audio recognition risk value and the image recognition risk value of the patient, respectively; The air leakage risk determination module is used to determine whether there is a risk of air leakage of the laryngeal mask according to the audio recognition risk value and the image recognition risk value.

9. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor and the memory communicate via the bus, and the processor executes the machine-readable instructions to perform the steps of any method as claimed in claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are executed.