Identification method, system and storage medium for Marsh grading
By combining the interactive guidance information of the image acquisition device, sound sensor and controller, audio and video information is collected in real time, and the correlation between acoustic features and airway physiological estimation features is used to screen target frames. This solves the problems of poor image acquisition quality and low coordination in the existing Mahalanobis grading method and achieves higher grading accuracy.
Patent Information
- Application Number
- CN202511006916.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-22
AI Technical Summary
The existing Mahalanobis grading method cannot fully expose the oral structure of the person being assessed, resulting in poor grading accuracy, mainly due to poor image acquisition quality and low cooperation of the person being assessed.
An image acquisition device, sound sensor and controller are used to collect audio and video information in real time by outputting interactive guidance information. The correlation between acoustic features and airway physiological estimation features is used to screen target frames, and the target frames are graded using the Mahalanobis hierarchical target recognition model.
It achieves high-quality image data acquisition, improves the accuracy of Mahalanobis grading, ensures that the oral structure of the person to be evaluated is fully exposed, and provides clearer and more effective image data to support accurate grading results.
Smart Images

Figure CN120511065B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of clinical medicine, and in particular to an identification method, system and storage medium for Mahalanobis grading. Background Art
[0002] With the continuous development of artificial intelligence technology, more and more difficult airway prediction models have been developed. These models are often based on facial recognition, combined with a modified Mahalanobis scale, and collect a variety of surface data. They have certain application value in predicting difficult airway. However, the existing modified Mahalanobis scale fails to fully expose the oral structure of the patient being evaluated, affecting the accuracy of the classification.
[0003] Typically, related technologies focus on facial image recognition. However, a relatively simple acquisition method is often used for facial image collection. For example, medical staff manually adjust the laryngoscope angle or directly use a camera to capture images, which is easily affected by subjective factors, resulting in poor image acquisition quality and recognition errors. Subsequently, image recognition algorithms are used to identify the photos and obtain Mahalanobis grading results. However, the degree of cooperation of the people to be evaluated is low, and they are often not clear about the movements required for Mahalanobis grading. In addition, the quality and effectiveness of the collected photos are difficult to guarantee, resulting in poor Mahalanobis grading recognition accuracy.
[0004] Therefore, how to improve the recognition accuracy of Mahalanobis grading has become a technical problem that needs to be solved urgently. Summary of the Invention
[0005] In view of this, the present invention provides a method, system and storage medium for identifying Mahalanobis scale, so as to solve the technical problem of how to improve the recognition accuracy of Mahalanobis scale.
[0006] In a first aspect, the present invention provides a method for identifying Mahalanobis scale, applicable to a system for identifying Mahalanobis scale. The identification system includes an image acquisition device, a sound sensor, and a controller. The identification method is applicable to the controller and includes:
[0007] Output interactive guidance information, wherein the interactive guidance information prompts the person to be identified to perform a detection action;
[0008] Acquiring audio information collected by the sound sensor when the person to be identified performs a detection action;
[0009] identifying a duration of target audio information in the audio information;
[0010] Collecting video information of the person to be identified performing a detection action during a duration of the target audio information;
[0011] Synchronously screening a target frame corresponding to the target audio information in time sequence in the video information based on the duration;
[0012] Target detection is performed on the target frame based on a Mahalanobis hierarchical target recognition model to obtain a Mahalanobis hierarchical recognition result.
[0013] In one embodiment, the step of synchronously screening the target frames corresponding to the valid audio information in time sequence in the video information based on the duration includes:
[0014] identifying acoustic features in the target audio information during the duration, and obtaining airway physiological estimation features based on the acoustic features;
[0015] identifying airway physiological image features of different video frames in the video information in a simultaneous sequence;
[0016] Calculating the correlation between the airway physiological estimation feature and the airway physiological image feature;
[0017] When the correlation is greater than the preset correlation, the current audio information is determined to be valid audio information, and the video frame corresponding to the valid audio information is determined to be the target frame.
[0018] In one embodiment, identifying acoustic features in the target audio information during the duration includes:
[0019] Extracting time domain features and frequency domain features from the target audio;
[0020] Determining the airflow velocity characteristics in the target audio based on pre-established mapping relationships between flow velocity and time domain characteristics and flow velocity and frequency domain characteristics;
[0021] Based on the correlation between airflow velocity and airway physiological structure, the degree of airway cross-sectional area expansion and mucosal brightness change corresponding to the airflow velocity characteristics are evaluated;
[0022] The degree of cross-sectional area expansion and the degree of mucosal brightness change are used as the airway physiological estimation features.
[0023] In one embodiment, identifying the duration of the target audio information in the audio information includes:
[0024] Use short-time energy and zero-crossing rate to detect the start and end time of audio information;
[0025] Extracting Mel-frequency cepstral coefficient features corresponding to the audio information within the start and end times;
[0026] The Mel-frequency cepstral coefficient feature is input into a pre-trained classification model for classification to obtain the target audio information and the duration corresponding to the target audio information.
[0027] In one embodiment, collecting video information of the person to be identified performing the detection action during the duration of the target audio information includes:
[0028] triggering the image acquisition device to acquire the video information when the target audio information is detected;
[0029] The video frames are screened using a non-maximum suppression algorithm to obtain a plurality of pre-selected video frames whose similarities are lower than a preset similarity.
[0030] In one embodiment, after obtaining a plurality of preselected video frames having similarities lower than a preset similarity, the method further includes:
[0031] Calculate the clarity, target space ratio and confidence of each pre-selected video frame;
[0032] Performing a weighted score on the clarity, target space ratio, and confidence level to obtain a validity score for each preselected video frame;
[0033] The video frames whose validity scores are greater than a preset score are selected as the video frames corresponding to the target audio information.
[0034] In one embodiment, performing target detection on the target frame based on the Mahalanobis hierarchical target recognition model to obtain a Mahalanobis hierarchical recognition result includes:
[0035] Extracting feature information of the target frame using a feature extraction network based on the ResNet50 network structure;
[0036] A classifier based on the random forest algorithm was used to score the severity of difficult airway using the extracted feature information, and the scoring results were mapped to Mahalanobis grading results.
[0037] In one embodiment, the detection action includes sticking out the tongue and exhaling, and the audio information includes exhalation audio information.
[0038] In a second aspect, the present invention provides an identification device for Mahalanobis classification, comprising:
[0039] A guidance module, configured to output interactive guidance information, wherein the interactive guidance information prompts the person to be identified to perform a detection action;
[0040] An acquisition module, configured to acquire audio information collected by the sound sensor when the person to be identified performs a detection action;
[0041] an identification module, configured to identify a duration of target audio information in the audio information;
[0042] An acquisition module, configured to acquire video information of the person to be identified performing a detection action during a duration of the target audio information;
[0043] A screening module, configured to synchronously screen target frames corresponding to the target audio information in time sequence in the video information based on the duration;
[0044] The detection module is used to perform target detection on the target frame based on the Mahalanobis hierarchical target recognition model to obtain a Mahalanobis hierarchical recognition result.
[0045] In a third aspect, the present invention provides a system for identifying Mahalanobis scales, comprising: an image acquisition device, a sound sensor, and a controller, wherein the controller comprises: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to thereby perform the method of the first aspect or any corresponding embodiment thereof.
[0046] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the method of the first aspect or any corresponding embodiment thereof.
[0047] The present invention provides a recognition method for Mahalanobis classification, which is applicable to a recognition system for Mahalanobis classification. The recognition system includes an image acquisition device, a sound sensor, and a controller. The recognition method is applicable to the controller and includes: outputting interactive guidance information, wherein the interactive guidance information prompts a person to be identified to perform a detection action; acquiring audio information collected by the sound sensor when the person to be identified performs the detection action; identifying a duration of target audio information in the audio information; acquiring video information of the person to be identified performing the detection action within the duration of the target audio information; synchronously screening target frames corresponding to the target audio information in time sequence in the video information based on the duration; and performing target detection on the target frames based on a Mahalanobis classification target recognition model to obtain a Mahalanobis classification recognition result. By using human-computer interaction, interactive guidance information is output to instruct the person to be evaluated to perform standard actions for Mahalanobis grading in real time. During the execution of the actions, audio information of the actions is collected in real time. The audio information is used to detect in real time whether the person to be evaluated has achieved the standard actions. When the standard actions are achieved, the image acquisition device is automatically triggered to collect video information with the standard actions and the duration of the target audio information. The target frames corresponding to the target audio (standard actions) in the video information are synchronously screened. This can achieve the acquisition of high-quality target frames, expose the oral structure of the person to be evaluated to the greatest extent, and provide clearer and more effective image data for Mahalanobis grading. The target recognition model for Mahalanobis grading is then used to perform target detection on the target frames to obtain accurate Mahalanobis grading recognition results.
[0048] Furthermore, in order to further capture the target frame corresponding to the valid audio (standard action), the acoustic features of the target audio information are identified during the duration, and the airway physiological estimation features are obtained based on the acoustic features. Simultaneously, the airway physiological image features of different video frames in the video information are respectively identified. The correlation between the airway physiological estimation features and the airway physiological image features is calculated. When the correlation is greater than a preset correlation, the current audio information is determined to be valid audio information, and the video frame corresponding to the valid audio information is determined to be the target frame. By utilizing the correlation between the airway physiological estimation features corresponding to the acoustic features of the vocalization and the actual captured image features, the optimal video frame corresponding to the standard action is selected, which can maximize the exposure of the oral structure of the person to be evaluated, provide clearer and more effective image data for Mahalanobis grading, and thus improve the accuracy of the grading. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0050] Figure 1 is a flow chart of a method for identifying Mahalanobis classification according to an embodiment of the present invention;
[0051] Figure 2 is a structural block diagram of a device for identifying Mahalanobis classification according to an embodiment of the present invention;
[0052] Figure 3 4 is a schematic diagram of the hardware structure of a controller in a Mahalanobis classification recognition system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0053] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.
[0054] According to an embodiment of the present invention, an embodiment of a method for identifying Mahalanobis scale is provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system, such as a set of computer-executable instructions. Moreover, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than that shown.
[0055] In one embodiment, a method for identifying Mahalanobis scales is applicable to a Mahalanobis scale identification system. The identification system includes an image acquisition device, an acoustic sensor, and a controller. The image acquisition device can be a high-speed camera, for example, one with a frame rate ≥120 fps and autofocus. The acoustic sensor can be a microphone with high sensitivity and wide bandwidth, for example, a sensitivity ≥-40 dB and a frequency response of 20 Hz-20 kHz. The identification system may also include a speaker and a display. In this embodiment, the identification method is applicable to the controller and includes:
[0056] S101. Output interactive guidance information, which prompts the person to be identified to perform a detection action. After the person to be identified is in place, interactive guidance information can be output via a speaker. This interactive guidance information can prompt the person to perform a detection action, where the detection action can include tongue extension and exhalation. It can also include actions such as head and neck position and mouth shape, for example, slightly tilting the head back 10-20 degrees, opening the mouth as wide as possible, and instructing the person to make a "ha" sound.
[0057] In this embodiment, standard movements, such as standard breathing movements, head tilting movements, tongue sticking out movements, etc., can also be displayed in real time on the display screen to provide a standard reference for the person to be evaluated.
[0058] In an optional embodiment, a virtual scene can be constructed using an MR device to present clear and intuitive guidance information to the person being evaluated. For example, the device can demonstrate the correct vocalization (exhalation) action in the form of a virtual animation, including requirements for exhalation force, duration, and rhythm, and further guide the person through voice prompts. The MR device tracks the person being evaluated's facial movements and posture in real time. When it detects that the person's facial posture is close to the preset vocalization preparation posture, it provides positive feedback to the person being evaluated, such as a flashing prompt icon in the virtual scene, encouraging the person to continue the vocalization operation. For example, after starting the virtual guidance program, the MR device presents a virtual mouth and airway model to the person being evaluated, accompanied by a voice prompt: "Please exhale slowly and steadily according to the animation on the screen." In the virtual animation, the model's mouth and airway simulate the correct exhalation action, demonstrating the airflow path and airway changes.
[0059] S102. Acquire audio information collected by the sound sensor while the person to be identified performs a detection action. In this embodiment, the sound sensor collects audio information of the person to be evaluated in real time while the person to be identified performs a detection action. For example, the sound of the person to be evaluated breathing out is collected in real time as audio information.
[0060] S103. Identify the duration of the target audio information within the audio information. In this embodiment, the audio information collected by the sound sensor may contain various noises, and other sound information may be collected before or after a breath. Therefore, in this embodiment, the target audio information can be identified within the collected audio information based on the characteristics corresponding to the breath sound. In this embodiment, Mel-frequency cepstral coefficient features can be extracted, and a classifier can be used to determine whether the breath is valid. The moment a valid breath is determined is used as the start time of the target audio, and the moment an invalid breath is determined is used as the end time, thereby determining the duration of the target audio information.
[0061] Exemplarily, the start and end times of audio information can be detected by using short-time energy and zero-crossing rate; extracting the Mel-frequency cepstral coefficient features corresponding to the audio information within the start and end times; inputting the Mel-frequency cepstral coefficient features into a pre-trained classification model for classification to obtain the target audio information and the duration corresponding to the target audio information.
[0062] Among them, the audio information can be preprocessed first. Wavelet denoising is used to remove environmental noises such as air conditioners and background human voices, and a soft threshold is set to filter the noise coefficient to improve the signal-to-noise ratio of the audio; the start and end times of exhalation are detected based on short-time energy and zero-crossing rate. Among them, the short-time energy STE and the zero-crossing rate ZCR are respectively:
[0063] ;
[0064] ;
[0065] Among them, x[n + k] represents the kth sample value starting from the nth sampling point of the audio information, which is used to intercept an audio information segment with a length of N. Here, n represents the starting position of the current analysis frame, and k represents the sampling point offset within the frame; w[k] is a Hamming window, N is the length of the Hamming window, and n is the frame index. By moving n, the short-time energy analysis of the signal is realized; sign() is a sign function that outputs the polarity of the signal, STE[n] is the short-time energy, and ZCR[n] is the zero-crossing rate.
[0066] In one embodiment, when the target audio (i.e., effective exhalation) is emitted, the airflow generates a sound with a specific frequency through the glottis, which has the characteristics of a sudden increase in short-time energy (STE increases) and a change in zero-crossing rate (ZCR first rises and then falls). Based on this, the thresholds of short-time energy STE[n] and zero-crossing rate ZCR[n] can be set to detect effective exhalation. Among them, when STE[n] > STE threshold and ZCR[n] < ZCR threshold, it is determined that exhalation starts. In order to avoid the influence of behaviors such as coughing and sighing, etc., the duration of STE[n] > STE threshold and ZCR[n] < ZCR threshold needs to exceed a preset time. For example, this preset time can be 200 ms, 500 ms, 1 s, etc., and it can be determined that effective exhalation starts.
[0067] Exemplarily, there is a large difference between the short-time energy during effective exhalation and the background noise energy. Therefore, the STE threshold can be set to 2 - 5 times the background noise energy. The zero-crossing rate of the exhalation sound is relatively low. Therefore, the ZCR threshold can be set to 50 - 100 times per second. If STE[n] > 2 - 5 times the background noise energy and ZCR[n] < 50 - 10 times per second, the corresponding exhalation is an effective exhalation, and the time period corresponding to continuous effective exhalations is used as the duration of the target audio information.
[0068] S104. Video information of the person to be identified performing the detection action is captured during the duration of the target audio information. When the target audio information is detected, a capture trigger signal is output to the image capture device at the start of the target audio to begin image capture. When the target audio stops, a stop trigger signal is output to the image capture device to stop triggering, thereby achieving synchronous audio and video capture. This improves the accuracy and specificity of video capture and reduces video storage overhead.
[0069] S105. Synchronously select target frames that are temporally corresponding to the target audio information in the video information based on the duration. In this embodiment, acoustic features are identified in the target audio information during the duration, and airway physiological estimation features are obtained based on the acoustic features, wherein the acoustic features are related to changes in the physiological structure of the airway. For example, the acoustic features may include airflow velocity during phonation, such as the airflow velocity during exhalation, and may also include power spectral density or Mel-frequency cepstral coefficients such as the main frequency and center of gravity frequency during phonation, wherein certain physiological changes occur in the corresponding vocal tract during exhalation.
[0070] When exhaling, the glottis opens, and the high-speed airflow passing through it triggers turbulence due to a sudden change in cross-section, generating broadband noise. The higher the airflow velocity, the greater the turbulence intensity and the higher the noise energy; the dominant frequency of turbulent noise is positively correlated with the flow velocity. Furthermore, at higher airflow velocities, the degree of expansion of the airway cross-sectional area and the degree of change in mucosal brightness are greater. For example, at airflow velocities greater than 5 m / s, the corresponding mucosal brightness in the video increases sharply, and the airway cross-sectional area expands by ≥15%. When the high-frequency portion of the sound information increases suddenly, the corresponding airway cross-sectional area and mucosal brightness also increase.
[0071] In one embodiment, the acoustic characteristics of a breath are calculated by extracting time-domain and frequency-domain features from the target audio. For example, a breath lacks periodic pulses and exhibits continuous broadband noise. The turbulence intensity of the airflow is determined by calculating the short-time energy, and / or the sound pressure is determined by calculating the root mean square of the noise energy. The turbulence intensity and / or sound pressure are used as time-domain features. Fast Fourier transform (FFT) is used to extract the dominant frequency or center of gravity frequency, or the Mel-frequency cepstral coefficients, as frequency-domain features.
[0072] The airflow velocity characteristics in the target audio are determined based on pre-established mapping relationships between flow velocity and time domain characteristics, and between flow velocity and frequency domain characteristics. In this embodiment, the airflow velocity characteristics in the target audio are calculated by mapping between flow velocity and time domain characteristics, and between flow velocity and frequency domain characteristics. In this embodiment, the mapping relationship between flow velocity and time domain characteristics can be: , where a and b are calibration coefficients and STE is short-time energy. The mapping relationship between flow velocity and frequency domain characteristics is: , where c and d are calibration coefficients, f c is the frequency domain feature.
[0073] In this embodiment, the calibration coefficients a, b, c, and d are obtained by fitting the actual acoustic characteristics (frequency domain characteristics and time domain characteristics) and the corresponding actual flow velocity. Specifically:
[0074] Multiple experimental or target individuals can be pre-guided to emit target audio signals of varying intensities (different airflow velocities), for example, by performing low, medium, and high-speed "blows." Simultaneously, the actual airflow velocity is collected using airflow velocity measurement equipment, such as a thermal airflow meter or a pressure gradient flowmeter. Sensors are also used to simultaneously record the original blow signal, and time and frequency domain features are extracted from the original signal. The time and frequency domain features are timestamped and aligned with the corresponding actual airflow velocity to construct corresponding sample pairs.
[0075] The empirical formula is a linear model: or , a linear regression algorithm is used to fit the sample pairs, solve the coefficients, and then obtain the empirical coefficients a, b, c, and d. Among them, the above-mentioned empirical coefficients can be solved using a linear regression algorithm such as the least squares method or the gradient descent method.
[0076] In an optional embodiment, a pre-trained support vector machine or neural network may be used to input features such as Mel-frequency cepstral coefficients, short-time energy, and main frequency, and output a flow rate value.
[0077] Based on the correlation between airflow velocity and airway physiological structure, for example, airflow velocity is positively correlated with the vocal tract cross-sectional area or vocal tract expansion speed and the mucosal brightness in the video. Therefore, after obtaining acoustic features such as airflow velocity, the degree of airway cross-sectional area expansion and the degree of mucosal brightness change corresponding to the airflow velocity features are evaluated; the degree of cross-sectional area expansion and the degree of mucosal brightness change are used as the airway physiological estimation features.
[0078] After obtaining the target audio information of an effective breath, the estimated airway physiological characteristics are evaluated using the acoustic characteristics of the effective breath. Simultaneously, the airway physiological image features of different video frames in the video information are identified. In this embodiment, after obtaining the video frame, the pharyngeal mucosa region in the video frame is segmented, the video frame is converted into a grayscale image, and the mean brightness of the pharyngeal mucosa region is calculated as the image mucosa brightness. Simultaneously, the edge contour of the pharyngeal mucosa region is extracted, and the edge contour area is calculated as the image acoustic tract cross-sectional area.
[0079] Calculate the correlation between the airway physiological estimation feature and the airway physiological image feature. Calculate the similarity between the mucosal brightness estimated using the acoustic feature and the image mucosal brightness, as well as the similarity between the vocal tract cross-sectional area estimated using the acoustic feature and the image vocal tract cross-sectional area, as the correlation between the airway physiological estimation feature and the airway physiological image feature. When the correlation is greater than the preset correlation, determine that the current audio information is valid audio information, and determine that the video frame corresponding to the valid audio information is the target frame. It can be considered that the image collected by the current person to be evaluated by performing the detection action can meet the image quality required by Mahalanobis grading, and the target frame can be obtained.
[0080] By identifying the acoustic features in the target audio information during the duration, and obtaining the airway physiological estimation features based on the acoustic features; simultaneously identifying the airway physiological image features of different video frames in the video information; calculating the correlation between the airway physiological estimation features and the airway physiological image features; when the correlation is greater than a preset correlation, determining that the current audio information is valid audio information, and determining that the video frame corresponding to the valid audio information is the target frame. By utilizing the correlation between the airway physiological estimation features corresponding to the acoustic features of the vocalization and the actual collected image features, the optimal video frame corresponding to the standard action is screened, which can maximize the exposure of the oral structure of the person to be evaluated, provide clearer and more effective image data for Mahalanobis grading, and thus improve the accuracy of subsequent Mahalanobis grading.
[0081] In this embodiment, the person to be evaluated is interactively guided to emit a standard target audio (such as an effective / standard exhalation action), which triggers the image acquisition device to start capturing video. By recognizing the target audio, it is possible to evaluate the airway physiological estimation features that the image acquisition device can theoretically capture when the person to be evaluated emits the target audio, and compare them with the airway physiological image features actually captured by the image acquisition device. When the similarity is greater than a preset similarity, it is considered that the person to be evaluated fully exposes the oral cavity or airway structure, and the optimal video frame is selected from the corresponding multiple video frames as the target frame to be graded.
[0082] S106. Perform target detection on the target frame based on the Mahalanobis graded target recognition model to obtain a Mahalanobis graded recognition result. Perform target detection and classification on the target frame based on the Mahalanobis graded target recognition model, output the target label and its confidence, and obtain the Mahalanobis graded result of the person to be evaluated. At the same time, perform data augmentation on the collected video data, including random cropping, horizontal and vertical flipping and other operations to augment the image data and improve the performance of the subsequent recognition algorithm. Specifically, a feature extraction network based on the ResNet50 network structure is used to extract the feature information of the target frame; a classifier constructed based on the random forest algorithm is used to score the severity of the difficult airway on the extracted feature information, and the scoring result is mapped to the Mahalanobis graded result.
[0083] In the present application, interactive guidance information is output, wherein the interactive guidance information prompts the person to be identified to perform a detection action; audio information collected by the sound sensor when the person to be identified performs the detection action is obtained; the duration of the target audio information in the audio information is identified; video information of the person to be identified when performing the detection action is collected within the duration of the target audio information; target frames corresponding to the target audio information in time sequence are synchronously screened in the video information based on the duration; target detection is performed on the target frame based on the Mahalanobis hierarchical target recognition model to obtain a Mahalanobis hierarchical recognition result. By utilizing human-computer interaction, interactive guidance information is output to instruct the person to be assessed to perform standard actions for Mahalanobis grading in real time. During the execution of the actions, audio information of the actions is collected in real time. The audio information is used to detect in real time whether the person to be assessed has achieved the standard actions. When the standard actions are achieved, the image acquisition device is automatically triggered to capture video information with the standard actions. During the duration of the target audio information, the target frames corresponding to the target audio (standard actions) in the video information are synchronously screened, which can achieve the acquisition of high-quality target frames, expose the oral structure of the person to be assessed to the greatest extent, and provide clearer and more effective image data for Mahalanobis grading. The target recognition model for Mahalanobis grading is then used to perform target detection on the target frames to obtain accurate Mahalanobis grading recognition results.
[0084] In one embodiment, if the target audio is identified (for example, the person to be evaluated emits a valid / standard breath), and the similarity between the airway physiological image features acquired by the corresponding image acquisition device and the airway physiological theoretical features evaluated by the target audio is greater than a preset similarity, it is also necessary to evaluate whether the airway physiological image features in the video frame meet the image features required for the Mahalanobis classification. If not, the current airway physiological image features are compared with the standard image features (the pre-set standard image features required for the Mahalanobis classification) to obtain the feature difference between the current airway physiological image features and the standard image features, and the interactive guidance information is updated based on the feature difference.
[0085] Specifically, updating the interactive guidance information based on the feature difference may include: determining the correlation between the feature difference and the detection action in the interactive guidance information, and the correlation may include the degree of difference, the degree of offset, etc. Exemplarily, for example, the feature difference may be the degree of exposure of the airway physiological structure of the current airway physiological image feature compared to the standard image feature, and the mouth shape adjustment action when exhaling is determined based on the exposure degree. For example, the exposure degree is 70%, and the corresponding mouth shape adjustment action when exhaling may be to increase the lip opening and closing by 30% before performing the exhaling action. And generate corresponding guidance information, for example, the display information of the lip opening and closing during the last exhalation action may be generated first, and the lip opening and closing corresponding to the generated guidance information may be superimposed on the display information to clearly instruct the person to be evaluated on how to adjust the action.
[0086] In this embodiment, when determining the difference features between the current airway physiological image features and the standard image features, the current airway physiological image can first be transformed according to the standard parameters of the standard image. For example, alignment transformations such as image stretching, translation, and scaling, as well as image quality transformations such as brightness and contrast, can be performed. After the transformation is completed, the transformed airway physiological image can be aligned with the standard image, and the two images can be subtracted to obtain the difference features. Furthermore, when subtracting the two images, key features can be extracted first. For example, the airway structural features in the two images can be extracted separately. After obtaining the airway structural features, the two airway structural features can be subtracted to obtain the difference features. The difference features are compared with the key features in the standard image to obtain the degree of difference.
[0087] In one embodiment, there are large individual differences between different persons to be evaluated. For example, special groups such as children and the elderly may have smaller short-term energy when exhaling; some other patients with airway diseases, such as asthma patients, have larger fluctuations in zero-crossing rate when exhaling. These individual differences may make it difficult to unify the breath detection threshold. Therefore, in this embodiment, the individual physiological characteristic baseline can be learned in real time and the detection threshold can be dynamically adjusted to ensure that persons to be evaluated of different ages and physical conditions can be accurately identified.
[0088] In this embodiment, the interactive guidance information may also include a respiratory feature baseline acquisition action, which is used to instruct the person to be evaluated to perform the respiratory feature baseline acquisition action, for example, prompting the person to be evaluated to perform 30 seconds to 1 minute of calm breathing. When the person to be evaluated performs the respiratory feature baseline acquisition action, the corresponding respiratory audio information and respiratory airflow rate are collected in real time, wherein the short-term energy mean STE during the calm breathing period is calculated based on the respiratory audio information. base and the mean zero crossing rate ZCR base , calculate and record the maximum airflow velocity V during quiet breathing by using the respiratory airflow velocity max and the average velocity fluctuation range Vbase , and obtain the respiratory characteristic baseline data.
[0089] Among them, the short-term energy mean STE base and the mean zero crossing rate ZCR base They are:
[0090] ;
[0091] ;
[0092] Where T is the number of frames corresponding to the baseline acquisition duration.
[0093] Calculate the standard deviation and coefficient of variation of baseline data to assess individual respiratory stability. The following formula can be used for calculation:
[0094] ;
[0095] Among them, σSTE is the standard deviation of short-time energy, CV STE is the short-time energy variation coefficient; σZCR is the standard deviation of the zero-crossing rate, CV ZCR is the coefficient of variation of the zero-crossing rate. In this embodiment, the standard deviation reflects the characteristic fluctuation range of normal breathing and is used to set the boundaries of the dynamic threshold. Individuals with a high coefficient of variation require a wider threshold tolerance. Based on this, the standard deviation and coefficient of variation of the baseline data can be used to set the threshold.
[0096] in:
[0097] STE threshold = (STE base +k1×σSTE)×(1+K×CV STE );
[0098] ZCR threshold = (ZCR base +k2×σZCR)×(1+K×CV ZCR );
[0099] Among them, k1 and k2 are adjustment coefficients, and K is the adaptive coefficient.
[0100] In this embodiment, the average flow velocity fluctuation range V base As the air flow velocity threshold.
[0101] The adaptive coefficient K is determined through cross-validation, and its value range is usually (0-1]; for normal adults, the normal mode k1=2.5 can be used, and for special groups such as the elderly and children, the preset low-intensity breathing mode k1=1.5 can be used; k1 is positively correlated with the breath intensity (i.e., airflow velocity) during normal breathing. And usually in the absence of respiratory diseases, k1≈k2. When there are respiratory diseases such as asthma that cause shortness of breath, k2 is usually set to a larger value, such as k2=3.0 or k2=3.2, and is set specifically according to the degree of shortness of breath. k2 is positively correlated with the degree of shortness of breath. The faster the breathing, the larger the value of k2.
[0102] In this embodiment, baseline data of calm breathing characteristics is collected for each individual in real time, and an adaptive threshold is set for each individual using the baseline data, so that target audio information (effective / standard breath) can be detected more accurately.
[0103] In one embodiment, before identifying the duration of the target audio information in the audio information, in order to ensure the diversity of video frames and avoid selecting similar frames, thereby reducing the overhead of computing resources and storage resources, in this embodiment, when the target audio information is detected, the image acquisition device is triggered to acquire the video information; and the video frames are screened using a non-maximum suppression algorithm to obtain a plurality of pre-selected video frames whose similarity is lower than a preset similarity.
[0104] In one embodiment, after obtaining the preselected video frames, it is also necessary to screen the quality of the video frames and filter out low-quality video frames to ensure recognition accuracy. In this embodiment, the clarity, target space proportion and confidence of each preselected video frame are calculated; the clarity, target space proportion and confidence are weighted and scored to obtain the effectiveness score of each preselected video frame; and the video frames with the effectiveness score greater than the preset score are screened as the video frames corresponding to the target audio information.
[0105] For example, a weighted comprehensive scoring mechanism is used to calculate the sharpness, target space ratio, and confidence of each frame image, and calculate the comprehensive score according to the weights.
[0106] Before performing the comprehensive scoring, clarity, target space ratio, and confidence can be dimensionless. Target space ratio and confidence are dimensionless data with a value range of 0-1. Clarity can be calculated based on the Laplace gradient variance and then dimensionlessly converted. Specifically, it can be calculated using the following formula:
[0107] ;
[0108] in, is the Laplace operator, which is used to detect the edges and details of the image; I(x,y) is the grayscale value of the pixel at the coordinate (x,y) of the preselected video frame; Var() is the variance calculation function, and the larger the variance, the higher the clarity.
[0109] In this embodiment, the clarity can be converted dimensionlessly by Min-Max normalization, that is, the clarity is converted into a numerical value in the interval [0,1] to adapt to the numerical range of the target space ratio and the confidence. Among them, the maximum value "Max" of the clarity and the minimum value "Min" of the clarity can be set in advance. For example, the theoretical minimum variance value and the theoretical maximum variance value can be used as the minimum value and the maximum value, or the actual boundary of the clarity fluctuation range in the video frames shot for multiple experimental individuals can be used as the minimum value and the maximum value. In the comprehensive scoring formula, the clarity is represented by the parameter "S", the target space ratio is represented by the parameter "TR", and the confidence is represented by the parameter "C". The following comprehensive scoring formula can be used for calculation.
[0110] The comprehensive scoring formula is:
[0111] Score = α × S + β × TR + γ × C;
[0112] Here, α is the clarity weight, β is the target space percentage weight, and γ is the confidence weight. The default weights are: α = 0.4, β = 0.3, and γ = 0.3. In airway assessment scenarios, image clarity determines the effectiveness of the information and is a prerequisite for identifying airway physiological characteristics. Effective identification is only possible when image clarity is sufficient. Therefore, when setting the default weights, the clarity weight can be set higher. For target space percentage and confidence, the default weights can be set using an average method.
[0113] Each weight can also be adjusted dynamically. For example, if a particular metric significantly exceeds a threshold (e.g., confidence > 0.9), its weight can be increased to ensure that high-quality frames are prioritized. If a majority of frames fail to meet a particular metric (e.g., target space percentage), an instruction to the user to adjust the weight can be output or the weight of that metric can be automatically reduced. In this embodiment, the clarity weight, target space percentage weight, and confidence weight must all be maintained within a range greater than 0 and less than 1 during the dynamic adjustment process.
[0114] This embodiment provides a recognition device for Mahalanobis classification, such as Figure 2 As shown, including:
[0115] The guidance module 201 is used to output interactive guidance information, wherein the interactive guidance information prompts the person to be identified to perform a detection action;
[0116] An acquisition module 202 is configured to acquire audio information collected by the sound sensor when the person to be identified performs a detection action;
[0117] An identification module 203 is configured to identify a duration of target audio information in the audio information;
[0118] The acquisition module 204 is configured to acquire video information of the person to be identified performing a detection action during a duration of the target audio information;
[0119] A screening module 205 is configured to synchronously screen target frames corresponding to the target audio information in time sequence in the video information based on the duration period;
[0120] The detection module 206 is configured to perform target detection on the target frame based on a Mahalanobis hierarchical target recognition model to obtain a Mahalanobis hierarchical recognition result.
[0121] It should be noted here that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiments.
[0122] It should be noted that the above modules as part of the device can be implemented through software or hardware, wherein the hardware environment includes a network environment.
[0123] An embodiment of the present invention further provides a system for identifying Mahalanobis scales, comprising an image acquisition device, a sound sensor, and a controller, wherein the controller comprises: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus, the memory is used to store a computer program, and the processor is used to execute the method in any of the above embodiments by running the computer program stored in the memory.
[0124] Figure 3 is a structural block diagram of an optional controller according to an embodiment of the present application, such as Figure 3 As shown, it includes a processor 10, a communication interface 20, a memory 30 and a communication bus 40, wherein the processor 10, the communication interface 20 and the memory 30 communicate with each other through the communication bus 40, wherein,
[0125] Memory 30, for storing computer programs;
[0126] The processor 10 is configured to execute the computer program stored in the memory 30, and implement the following method:
[0127] Output interactive guidance information, wherein the interactive guidance information prompts the person to be identified to perform a detection action;
[0128] Acquiring audio information collected by the sound sensor when the person to be identified performs a detection action;
[0129] identifying a duration of target audio information in the audio information;
[0130] Collecting video information of the person to be identified performing a detection action during a duration of the target audio information;
[0131] Synchronously screening a target frame corresponding to the target audio information in time sequence in the video information based on the duration;
[0132] Target detection is performed on the target frame based on a Mahalanobis hierarchical target recognition model to obtain a Mahalanobis hierarchical recognition result.
[0133] Optionally, in this embodiment, the communication bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The communication bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0134] The communication interface is used for communication between the above controller and other devices.
[0135] The memory may include RAM, or may include non-volatile memory, such as at least one disk memory. Alternatively, the memory may also be at least one storage device located away from the aforementioned processor.
[0136] The above-mentioned processor can be a general-purpose processor, which can include but is not limited to: CPU (Central Processing Unit), NP (Network Processor), etc.; it can also be DSP (Digital Signal Processing), ASIC (Application Specific Integrated Circuit), FPGA (Field-Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0137] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, and this embodiment will not be described in detail here.
[0138] It can be understood by those skilled in the art that Figure 3 The structure shown is for illustration only. The device for implementing any one of the methods in the above embodiments may be a terminal device, which may be a smart phone (such as an Android phone, an IOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 3 It does not limit the structure of the above electronic device. For example, the terminal device may also include Figure 3 More or fewer components (such as network interfaces, display devices, etc.) shown in, or with Figure 3 Different configurations shown.
[0139] A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, which can include: a flash drive, ROM, RAM, a magnetic disk or an optical disk, etc.
[0140] As an exemplary embodiment, the present application further provides a computer-readable storage medium, wherein the storage medium stores a computer program, wherein the computer program is configured to execute the following method steps when executed:
[0141] Output interactive guidance information, wherein the interactive guidance information prompts the person to be identified to perform a detection action;
[0142] Acquiring audio information collected by the sound sensor when the person to be identified performs a detection action;
[0143] identifying a duration of target audio information in the audio information;
[0144] Collecting video information of the person to be identified performing a detection action during a duration of the target audio information;
[0145] Synchronously screening a target frame corresponding to the target audio information in time sequence in the video information based on the duration;
[0146] Target detection is performed on the target frame based on a Mahalanobis hierarchical target recognition model to obtain a Mahalanobis hierarchical recognition result.
[0147] Optionally, in this embodiment, the above-mentioned storage medium can be used to execute the program code of the method steps of the embodiment of the present application.
[0148] Optionally, in this embodiment, the above-mentioned storage medium may be located on at least one network device among the multiple network devices in the network shown in the above-mentioned embodiment.
[0149] Optionally, in this embodiment, the storage medium is configured to store data for executing the method in the above embodiment.
[0150] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments, which will not be described in detail in this embodiment.
[0151] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media that can store program codes, such as a USB flash drive, a ROM, a RAM, a mobile hard disk, a magnetic disk, or an optical disk.
[0152] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0153] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing one or more computer devices (such as personal computers, servers, or network devices) to execute all or part of the steps of the method in the above embodiments.
[0154] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, there may be other division methods, such as combining or integrating multiple units or components into another system, or ignoring or not implementing some features. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interface, indirect coupling or communication connection of units or modules, and may be electrical or other forms.
[0155] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the purpose of the solution provided in this embodiment.
[0156] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0157] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0158] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for identifying Mahalanobis classification, characterized in that: A recognition system applicable to Mahalanobis grading includes an image acquisition device, a sound sensor, and a controller. The recognition method is applicable to the controller and includes: Output interactive guidance information, wherein the interactive guidance information prompts the person to be identified to perform a detection action; Acquiring audio information collected by the sound sensor when the person to be identified performs a detection action; identifying a duration of target audio information in the audio information; Collecting video information of the person to be identified performing a detection action during a duration of the target audio information; Synchronously screening a target frame corresponding to the target audio information in time sequence in the video information based on the duration period, comprising: identifying acoustic features in the target audio information during the duration, and obtaining airway physiological estimation features based on the acoustic features; identifying airway physiological image features of different video frames in the video information in a simultaneous sequence; Calculating the correlation between the airway physiological estimation feature and the airway physiological image feature; When the correlation is greater than a preset correlation, determining that the current audio information is valid audio information, and determining that the video frame corresponding to the valid audio information is the target frame; The step of identifying the acoustic features in the target audio information during the duration period includes: Extracting time domain features and frequency domain features from the target audio; Determining the airflow velocity characteristics in the target audio based on pre-established mapping relationships between flow velocity and time domain characteristics and flow velocity and frequency domain characteristics; Based on the correlation between airflow velocity and airway physiological structure, the degree of airway cross-sectional area expansion and mucosal brightness change corresponding to the airflow velocity characteristics are evaluated; The cross-sectional area expansion degree and the mucosal brightness change degree are used as the airway physiological estimation features; Target detection is performed on the target frame based on a Mahalanobis hierarchical target recognition model to obtain a Mahalanobis hierarchical recognition result.
2. The method for identifying Mahalanobis classification according to claim 1, wherein: The identifying a duration of target audio information in the audio information comprises: Use short-time energy and zero-crossing rate to detect the start and end time of audio information; Extracting Mel-frequency cepstral coefficient features corresponding to the audio information within the start and end times; The Mel-frequency cepstral coefficient feature is input into a pre-trained classification model for classification to obtain the target audio information and the duration corresponding to the target audio information.
3. The identification method for Mahalanobis classification according to claim 1, characterized in that: The collecting of video information of the person to be identified performing the detection action during the duration of the target audio information includes: triggering the image acquisition device to acquire the video information when the target audio information is detected; The video frames are screened using a non-maximum suppression algorithm to obtain a plurality of pre-selected video frames whose similarities are lower than a preset similarity.
4. The identification method for Mahalanobis classification according to claim 3, characterized in that: After obtaining a plurality of preselected video frames whose similarities are lower than a preset similarity, the method further includes: Calculate the clarity, target space ratio and confidence of each pre-selected video frame; Performing a weighted score on the clarity, target space ratio, and confidence level to obtain a validity score for each preselected video frame; The video frames whose validity scores are greater than a preset score are selected as the video frames corresponding to the target audio information.
5. The identification method for Mahalanobis classification according to claim 1, wherein: The target detection is performed on the target frame based on the Mahalanobis hierarchical target recognition model to obtain the Mahalanobis hierarchical recognition result, which includes: Extracting feature information of the target frame using a feature extraction network based on the ResNet50 network structure; A classifier based on the random forest algorithm was used to score the severity of difficult airway using the extracted feature information, and the scoring results were mapped to Mahalanobis grading results.
6. The method for identifying Mahalanobis classification according to any one of claims 1 to 5, characterized in that: The detection actions include sticking out the tongue and breathing out, and the audio information includes breathing out audio information.
7. A recognition system for Mahalanobis classification, characterized in that: include: An image acquisition device, a sound sensor, and a controller, wherein the controller comprises: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the identification method for Mahalanobis classification according to any one of claims 1 to 6 by executing the computer instructions.
8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the identification method for Mahalanobis classification according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-mode flexible tracheal intubation system and position identification method
CN116966382A
Trachea cannula guiding method and guiding type laryngoscope
CN120022488A