A pathological voice information processing method and system

By constructing a nasopharyngeal dimension matrix and analyzing pathological voice information, combined with doctors' experience, the problem of low accuracy in pathological voice analysis was solved, enabling more accurate early assessment.

CN121096381BActive Publication Date: 2026-04-10THE SECOND AFFILIATED HOSPITAL TO NANCHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-10
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

The accuracy of analysis results for pathological voices in existing technologies is low, resulting in low reference value for early auxiliary assessment results of patients' pathological voices.

Method used

By constructing a nasopharyngeal dimension matrix, obtaining dimension-related feature vectors and feature values, analyzing abnormal fluctuations in dimension-related directional components, and combining doctors' clinical experience, the pathological value of speech is integrated, and the pathological voice degree is quantified.

Benefits of technology

It improves the accuracy of patient voice information recognition and enhances the reference value of early auxiliary assessment results of pathological voices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121096381B_ABST
    Figure CN121096381B_ABST
Patent Text Reader

Abstract

The application discloses a pathological voice information processing method and system, and relates to the technical field of speech processing. The method comprises the following steps: constructing a nasopharyngeal dimension matrix through a nasopharyngeal sound segment according to a plurality of dimension characteristic values in different dimensions; then, a plurality of dimension-related direction components are obtained; abnormal fluctuations of data on the dimension-related direction components are analyzed, and the dimension-related characteristic values are combined to serve as speech pathology values of the dimension-related direction components; the values of abnormal fluctuations on the dimension-related direction components are analyzed, and speech pathology values of different dimension-related direction components are integrated to obtain pathological voice degrees of phoneme types; and then, early auxiliary evaluation is carried out. The application improves the accuracy of patient voice information recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech processing, in particular to a pathological voice information processing method and system. BACKGROUND

[0002] Voice is one of the important physiological characteristics of human body, and the generation of normal voice depends on complex physiological mechanisms, including lung air flow causing regular vibration of vocal cords, generating specific characteristic sound waves, and forming after oral-nasal pharynx resonance. When the throat is diseased, the acoustic parameters of pathological voice signals will deviate from the normal state. Therefore, by analyzing the pathological voice of patients, important clues can be provided for early assessment of diseases.

[0003] The prior art usually only analyzes from the acoustic angle, and does not effectively combine with the doctor's inquiry experience, resulting in low accuracy of the analysis results of the abnormal conditions of the patient's voice, and low value of the early auxiliary assessment results of the patient's pathological voice for the doctor's reference. SUMMARY

[0004] Therefore, the purpose of the present application is to provide a pathological voice information processing method and system, which aims to solve the problem of low accuracy of the analysis results of the pathological voice in the prior art, and low value of the early auxiliary assessment results of the patient's pathological voice for the doctor's reference.

[0005] The present application provides a pathological voice information processing method, which comprises:

[0006] A pathological voice information processing method, characterized in that the method comprises:

[0007] Obtaining a plurality of nasopharyngeal sound segments of the same phoneme category in different dimensions, wherein the nasopharyngeal sound segments correspond to a dimension feature value in different dimensions respectively;

[0008] According to a plurality of dimension feature values in different dimensions, a nasopharyngeal dimension matrix is constructed by using the nasopharyngeal sound segments; feature analysis is performed on the nasopharyngeal dimension matrix to obtain a dimension-related feature vector and a dimension-related feature value; and according to the dimension-related feature vector, a dimension-related direction component in the direction indicated by each dimension-related feature vector is obtained from the nasopharyngeal dimension matrix;

[0009] Abnormal fluctuations of data on the dimension-related direction component are analyzed as a pathological performance degree of each dimension-related direction component; the value obtained by combining the dimension-related feature value and the pathological performance degree is taken as the voice pathology value of each dimension-related direction component; the values of abnormal fluctuations on the dimension-related direction components are analyzed, and the value obtained by integrating the voice pathology values of different dimension-related direction components is taken as the pathological voice degree of the phoneme category.

[0010] Preferably, the method for obtaining the nasopharyngeal dimension matrix comprises the following steps:

[0011] For any one dimension, sequence construction is performed on the dimension characteristic values of the dimension on all nasopharyngeal phonemes, and the constructed sequence is taken as the speech dimension representation sequence of the dimension.

[0012] The speech dimension representation sequence of each dimension is obtained.

[0013] The speech dimension representation sequence of each dimension is taken as a column vector, and matrix construction is performed, and the constructed matrix is taken as the nasopharyngeal dimension matrix.

[0014] Preferably, the method for obtaining the dimension-related characteristic vector and the dimension-related characteristic value comprises the following steps:

[0015] Eigenvalue decomposition is performed on the nasopharyngeal dimension matrix to obtain a plurality of characteristic vectors and a plurality of characteristic values, each characteristic vector is taken as a dimension-related characteristic vector, and each characteristic value is taken as a dimension-related characteristic value; each dimension-related characteristic vector corresponds to a dimension-related characteristic value.

[0016] Preferably, the method for obtaining the dimension-related directional component comprises the following steps:

[0017] For any one dimension-related characteristic vector, a vector obtained by multiplying the dimension-related characteristic vector with the nasopharyngeal dimension matrix is taken as a dimension-related directional component in the direction indicated by the dimension-related characteristic vector.

[0018] Preferably, the method for obtaining the pathological manifestation degree comprises the following steps:

[0019] For any one dimension-related directional component, a plurality of component abnormal data of the dimension-related directional component are obtained.

[0020] The overall average amplitude of all component abnormal data is taken as the pathological manifestation degree of the dimension-related directional component.

[0021] Preferably, the method for obtaining the component abnormal data comprises the following steps:

[0022] Anomaly detection is performed on the data on the dimension-related directional component by using an anomaly detection algorithm, a plurality of abnormal data of the dimension-related directional component are obtained, and each abnormal data is taken as component abnormal data.

[0023] Preferably, the numerical value obtained by combining the dimension-related characteristic value with the pathological manifestation degree is taken as the speech pathological value of each dimension-related directional component; the numerical value obtained by analyzing the abnormal fluctuation of the dimension-related directional component and integrating the speech pathological values of different dimension-related directional components is taken as the pathological voice degree of the phoneme category.

[0024] For any one dimension-related direction component, the value combined by the dimension characteristic value on the dimension-related characteristic vector to which the dimension-related direction component belongs and the pathological manifestation degree is taken as the speech pathology value of the dimension-related direction component;

[0025] For any one dimension-related direction component, the component pathological voice degree of the dimension-related direction component is obtained according to the pathological manifestation degree and the speech pathology value;

[0026] The component pathological voice degree of each dimension-related direction component is obtained.

[0027] The cumulative mapping result of the component pathological voice degrees of all the dimension-related direction components is taken as the pathological voice degree of the phoneme category.

[0028] The value combined by the pathological manifestation degree and the speech pathology value of the dimension-related direction component is taken as the component pathological voice degree of the dimension-related direction component.

[0029] Preferably, the method for obtaining the component pathological voice degree comprises:

[0030] The value combined by the pathological manifestation degree and the speech pathology value of the dimension-related direction component is taken as the component pathological voice degree of the dimension-related direction component.

[0031] Preferably, after the step of analyzing the abnormal fluctuation values of the dimension-related direction components and integrating the speech pathology values of different dimension-related direction components to obtain the pathological voice degree of the phoneme category, the system further comprises:

[0032] The pathological voice degree is normalized.

[0033] Another aspect of the present application provides a pathological voice information processing system, which comprises:

[0034] A nasopharyngeal sound segment acquisition module is configured to acquire a plurality of nasopharyngeal sound segments of the same phoneme category under different dimensions, wherein each nasopharyngeal sound segment corresponds to a dimension characteristic value under a different dimension.

[0035] A dimension-related direction component acquisition module is configured to construct a nasopharyngeal dimension matrix by using the nasopharyngeal sound segments according to the plurality of dimension characteristic values under different dimensions; perform feature analysis on the nasopharyngeal dimension matrix to obtain a dimension-related characteristic vector and a dimension-related characteristic value; and obtain a dimension-related direction component in the direction indicated by each dimension-related characteristic vector from the nasopharyngeal dimension matrix.

[0036] The pathological voice degree acquisition module is used for analyzing abnormal fluctuations of data on the dimension-related directional component as a pathological performance degree of each dimension-related directional component; combining the dimension-related characteristic value with the pathological performance degree to obtain a numerical value of each dimension-related directional component as a voice pathological value of each dimension-related directional component; analyzing the numerical value of abnormal fluctuations on the dimension-related directional component, and integrating the voice pathological values of different dimension-related directional components to obtain a pathological voice degree of the phoneme category.

[0037] According to the dimension characteristic values in different dimensions, the nasopharynx dimension matrix is constructed, and the dimension-related directional components are obtained from the nasopharynx dimension matrix; the dimension-related directional components are used for describing component information associated with audio information mainly concerned by doctors in each dimension, and the experience of doctors in medical treatment is integrated into each dimension of the detected voice information, so that the association between each dimension of the voice information is more intelligent; then the voice pathological values of the dimension-related directional components are obtained by analyzing abnormal fluctuations of data on the dimension-related directional components and combining the dimension-related characteristic values; then the numerical value of abnormal fluctuations on the dimension-related directional component is analyzed, and the voice pathological values of different dimension-related directional components are integrated to obtain a pathological voice degree of the phoneme category; the pathological voice degree is used for describing the obvious degree of pathological characteristics of the patient's pronunciation part when pronouncing the corresponding phoneme category, and can quantize the pathological relationship of the patient when pronouncing different phoneme categories; the accuracy of voice information recognition of the patient is improved, the early auxiliary evaluation result of the pathological voice of the patient can be used as a reference for the doctor, and the problem that the accuracy of the analysis result of the pathological voice in the prior art is low and the early auxiliary evaluation result of the pathological voice of the patient can be used as a reference for the doctor is solved. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 The flow chart of the pathological voice information processing method in the first embodiment of the present application is shown in the figure.

[0039] Figure 2 The structural block diagram of the pathological voice information processing system in the second embodiment of the present application is shown in the figure.

[0040] The following specific embodiments will further illustrate the present application in combination with the above-mentioned figures. DETAILED DESCRIPTION

[0041] In order to facilitate the understanding of the present application, the present application will be described more fully below with reference to the related drawings. The figures show several embodiments of the present application. However, the present application can be realized in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0042] It should be noted that when an element is referred to as being "on" another element, it can be directly on the other element or intervening elements can also be present. When an element is referred to as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element or intervening elements can also be present. The terms "vertical", "horizontal", "left", "right", and the like as used herein are used for illustration only and do not limit the scope of the application.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0044] Embodiment one

[0045] Referring to Figure 1 , a method for processing pathological voice information in a first embodiment of the application is shown, and the method comprises steps S10-S12.

[0046] In step S10, a plurality of nasopharyngeal sound segments of the same phoneme category in different dimensions are obtained, and the nasopharyngeal sound segments correspond to a dimension feature value in different dimensions respectively.

[0047] It should be noted that the prior art usually only analyzes from the acoustic angle, and does not effectively combine with the doctor's experience in inquiring, resulting in that the analysis result of the abnormality of the patient's voice onset is low in precision, and the early auxiliary evaluation result of the patient's pathological voice has low value for the doctor to refer to.

[0048] In the embodiment of the present application, the acquisition of the nasopharyngeal sound segment is to collect and preprocess the original speech data to provide standardized input for subsequent analysis, focusing on the feature differences of the same phoneme in different dimensions. In one specific implementation of the embodiment of the present application, the acquisition process of the nasopharyngeal sound segment is as follows: acquiring a voice audio detection sequence of any patient from a patient voice database; regarding each speech data segment that needs to be focused on in the voice audio detection sequence as a nasopharyngeal audio data segment; taking any one nasopharyngeal audio data segment as an example, acquiring a plurality of phoneme categories in the nasopharyngeal audio data segment through a trained neural network; wherein the neural network used in the embodiment is RNN, and the data set for training the neural network is acquired by: collecting a large amount of nasopharyngeal audio data segments, manually marking the phoneme categories in each nasopharyngeal audio data segment as a marking result, and taking the marking result as the label of each nasopharyngeal audio data segment; collecting a large amount of nasopharyngeal audio data segments and corresponding labels to form a data set; training the neural network using the data set, and the loss function used in the training process is RNN-T loss function; wherein the specific training process is a known content of the neural network, and the embodiment will not repeat the specific training process. Acquire a plurality of phoneme categories in each nasopharyngeal audio data segment. Each phoneme category corresponds to a plurality of audio data in each nasopharyngeal audio data segment.

[0049] In specific implementation, the phonemes highly related to nasopharyngeal lesions in clinical practice (such as nasal sounds / half-nasal sounds "m", "n", "ng", etc.) can be selected because the pronunciation of such phonemes depends on the resonance of the nasopharynx, and the characteristic changes are more significant when the lesion occurs. For each phoneme pronunciation segment (nasopharyngeal sound segment), a plurality of acoustic dimension features can be extracted.

[0050] It should be noted that the patient voice database in the embodiment stores a plurality of voice detection data sequences of patients, wherein the voice detection data sequence is an audio data sequence, and each voice detection data sequence contains a plurality of speech data segments that need to be focused on and are marked by doctors through clinical experience.

[0051] Further, taking any one phoneme category as an example, a nasopharyngeal audio data segment containing the phoneme category is taken as a phoneme category data segment of the phoneme category; each audio data corresponding to the phoneme category in the phoneme category data segment is taken as data under a frame, Mel frequency cepstrum coefficient (MFCC) conversion is performed, a Mel frequency cepstrum coefficient sequence of the phoneme category on the corresponding data is obtained, and the Mel frequency cepstrum coefficient sequence is taken as a nasopharyngeal audio data segment of the phoneme category in the phoneme category data segment; all nasopharyngeal audio data segments of the phoneme category in the phoneme category data segment are obtained. The number of data corresponding to each phoneme category in the corresponding phoneme category data segment is not uniform; each data corresponding to each phoneme category in the corresponding phoneme category data segment corresponds to a Mel frequency cepstrum coefficient sequence, each Mel frequency cepstrum coefficient sequence contains 13 coefficients, and each coefficient represents a dimension.

[0052] It is particularly pointed out that the process of obtaining a Mel frequency cepstrum coefficient sequence according to audio data under different frames is a known technology, and the embodiment will not be described again.

[0053] Further, in all nasopharyngeal audio data segments on the phoneme category data segment, the coefficients under the same sequence number in the nasopharyngeal audio data segment are sorted in the order of time from early to late, and the data segment formed after the sorting is taken as a nasopharyngeal audio segment of the phoneme category under the same dimension; the nasopharyngeal audio segment is subjected to Fourier transform to obtain a frequency spectrum graph, and the frequency with the largest amplitude in the frequency spectrum graph is taken as a dimension characteristic value of the nasopharyngeal audio segment. A plurality of nasopharyngeal audio segments of the phoneme category under the same dimension and corresponding dimension characteristic values are obtained. The number of nasopharyngeal audio segments of the phoneme category under the same dimension is consistent with the number of phoneme category data segments of the phoneme category.

[0054] It is particularly pointed out that the process of Fourier transform to obtain a frequency spectrum graph is a known technology, and the embodiment will not be described again.

[0055] At this point, a plurality of nasopharyngeal audio segments of the phoneme category under different dimensions are obtained by the above method.

[0056] In step S11, a nasopharyngeal dimension matrix is constructed by using the nasopharyngeal audio segments according to a plurality of dimension characteristic values under different dimensions; feature analysis is performed on the nasopharyngeal dimension matrix to obtain a dimension-related feature vector and a dimension-related feature value; and a dimension-related direction component in the direction indicated by each dimension-related feature vector is obtained from the nasopharyngeal dimension matrix according to the dimension-related feature vector.

[0057] It should be noted that different sound emission in the acoustic system causes different acoustic feature changes, and in order to facilitate the doctor to judge the disease of the patient's voice, the patient usually reads the specified text content for detection, and the text content itself has obvious pronunciation rules, so that the voice produced by the patient after reading will also have certain correlation in each dimension. The greater the correlation between certain dimensions, the more complex the pronunciation of the audio content in the corresponding dimension, the more valuable it is. Therefore, according to the dimension characteristic values in different dimensions, the nasopharyngeal dimension matrix can be constructed by using the nasopharyngeal phonetic segment; the dimension correlation feature vector and the dimension correlation feature value are obtained by performing feature analysis on the nasopharyngeal dimension matrix; and the dimension correlation direction component in the direction indicated by each dimension correlation feature vector is obtained from the nasopharyngeal dimension matrix according to the dimension correlation feature vector.

[0058] Preferably, in some implementations of the embodiments of the present application, the method for obtaining the nasopharyngeal dimension matrix is: for any dimension, the dimension characteristic values of the dimension on all nasopharyngeal phonetic segments are sequentially constructed, and the constructed sequence is taken as the speech dimension representation sequence of the dimension; the speech dimension representation sequence of each dimension is obtained; each speech dimension representation sequence is taken as a column vector, and a matrix is constructed, and the constructed matrix is taken as the nasopharyngeal dimension matrix. The specific process is as follows:

[0059] Taking any one dimension of the phoneme category as an example, the sequence of the dimension characteristic values of all nasopharyngeal phonetic segments of the phoneme category in the dimension is taken as the speech dimension representation sequence of the phoneme category in the dimension; the speech dimension representation sequence of the phoneme category in each dimension is obtained; the first element in each speech dimension representation sequence is taken as the first element in the column vector, and the last element is taken as the last element in the column vector, a column vector is constructed, and a covariance matrix is constructed according to the column vector, and the constructed covariance matrix is taken as the nasopharyngeal dimension matrix of the phoneme category. The process of obtaining the covariance matrix according to the column vector is a known technology, and will not be described herein.

[0060] Preferably, in some implementations of the embodiments of the present application, the method for obtaining the dimension correlation feature vector and the dimension correlation feature value is: performing eigenvalue decomposition on the nasopharyngeal dimension matrix to obtain a plurality of characteristic vectors and a plurality of characteristic values, taking each characteristic vector as a dimension correlation feature vector, and taking each characteristic value as a dimension correlation feature value; each dimension correlation feature vector corresponds to a dimension correlation feature value. The specific process is as follows:

[0061] Eigenvalue decomposition is performed on the nasopharyngeal dimension matrix of the phoneme category to obtain a plurality of characteristic vectors and a plurality of characteristic values; each characteristic vector is taken as a dimension correlation feature vector, and each characteristic value is taken as a dimension correlation feature value. Each dimension correlation feature vector corresponds to a dimension correlation feature value.

[0062] Preferably, in some implementations of the embodiments of the present application, the method for obtaining the dimension-related direction component is: for any one dimension-related feature vector, the vector obtained after multiplying the dimension-related feature vector with the nasopharyngeal dimension matrix is taken as the dimension-related direction component in the direction indicated by the dimension-related feature vector. The specific process is as follows:

[0063] Taking any one dimension-related feature vector as an example, the dimension-related feature vector is multiplied with the nasopharyngeal dimension matrix of the phoneme category to obtain a new vector, which is taken as the dimension-related direction component in the direction indicated by the dimension-related feature vector. The dimension-related direction component in the direction indicated by each dimension-related feature vector is obtained.

[0064] Thus far, the dimension-related direction component in the direction indicated by each dimension-related feature vector is obtained by the above method.

[0065] In step S12, the abnormal fluctuation of data on the dimension-related direction component is analyzed as the pathological performance degree of each dimension-related direction component; the value obtained by combining the dimension-related feature value with the pathological performance degree is taken as the speech pathology value of each dimension-related direction component; the value obtained by analyzing the abnormal fluctuation of data on the dimension-related direction component and integrating the speech pathology values of different dimension-related direction components is taken as the pathological voice degree of the phoneme category.

[0066] It should be noted that if the patient to be detected has a corresponding disease, the patient will have irregular respiratory movement, which will affect the stability of the lung airflow, cause the vocal cords to bend and incompletely close, and thus cause the acoustic characteristics of the patient to change when the patient pronounces a sound that requires the vocal cords to close. At the same time, because different phonemes have different performances when pronouncing, the acoustic characteristics of the voice will also change in normal pronunciation data, but there is still a significant difference compared with the acoustic characteristic change caused by pathological voice. Therefore, the abnormal fluctuation of data on the dimension-related direction component can be analyzed as the pathological performance degree of each dimension-related direction component; the value obtained by combining the dimension-related feature value with the pathological performance degree is taken as the speech pathology value of each dimension-related direction component; the value obtained by analyzing the abnormal fluctuation of data on the dimension-related direction component and integrating the speech pathology values of different dimension-related direction components is taken as the pathological voice degree of the phoneme category.

[0067] Preferably, in some implementations of the embodiments of the present application, the method for obtaining the pathological performance degree is: for any one dimension-related direction component, a plurality of component abnormal data of the dimension-related direction component are obtained; the overall average amplitude of all component abnormal data is taken as the pathological performance degree of the dimension-related direction component. The specific process is as follows:

[0068] Preferably, in some implementations of the present invention, the method for obtaining component anomaly data is as follows: anomaly detection algorithms are used to detect anomalies in the data of dimension-related direction components, obtaining several anomaly data points for the dimension-related direction components, and each anomaly data point is taken as component anomaly data. The specific process is as follows:

[0069] Taking any one-dimensional related directional component as an example, through The anomaly detection algorithm performs anomaly detection on all data in the relevant directional components of this dimension, obtaining several anomalous data points; and treats each anomalous data point as a component anomaly. The process of anomaly detection to obtain anomalous data is as follows: The well-known content of anomaly detection algorithms will not be repeated in this embodiment.

[0070] Furthermore, the mean of all abnormal component data is used as the pathological manifestation degree of the relevant directional component in that dimension.

[0071] Preferably, in some implementations of the present invention, the method for obtaining the speech pathology value is as follows: for any dimension-related directional component, the value obtained by combining the dimensional feature value of the dimension-related feature vector to which the dimension-related directional component belongs with the pathological manifestation degree is taken as the speech pathology value of the dimension-related directional component. The specific process is as follows:

[0072] The product of the dimensional eigenvalue of the dimensional directional component and the pathological manifestation degree is taken as the speech pathological value of the dimensional directional component.

[0073] Preferably, in some implementations of the present invention, the method for obtaining pathological voice intensity is as follows: for any dimension-related directional component, based on the pathological manifestation degree and the speech pathology value, the component pathological voice intensity of the dimension-related directional component is obtained; the component pathological voice intensity of each dimension-related directional component is obtained; the cumulative mapping result of the component pathological voice intensity of all dimension-related directional components is used as the pathological voice intensity of the phoneme type. The specific process is as follows:

[0074] Preferably, in some implementations of the present invention, the method for obtaining the component pathological voice level is as follows: the value obtained by combining the pathological manifestation degree of the dimension-related directional component with the speech pathological value is used as the component pathological voice level of the dimension-related directional component. The specific process is as follows:

[0075] The product of the pathological manifestation degree of the relevant directional component of this dimension and the pathological value of the speech is taken as the component pathological voice degree of the relevant directional component of this dimension.

[0076] Furthermore, the component pathological voice intensity of each dimension-related directional component is obtained; the normalized value of the sum of the component pathological voice intensity of all dimensions-related directional components is used as the pathological voice intensity of that phoneme type.

[0077] It should be noted that the greater the pathological vocal intensity, the more likely there is an abnormality in the corresponding phoneme type, and the more obvious the pathological characteristics are when the patient's vocalization site produces a sound of that phoneme type.

[0078] It should be noted that this embodiment defaults to... The function is normalized, and the normalization function can be determined according to the specific implementation. This embodiment will not elaborate on it.

[0079] Thus, the pathological vocal quality of phoneme types was obtained through the above method.

[0080] In practice, early auxiliary assessments should be conducted based on pathological vocal quality.

[0081] Preferably, in a specific implementation of this invention, the specific process of early auxiliary assessment based on pathological noise level is as follows: a pathological noise level threshold is preset. The pathological voice intensity is greater than The phoneme types are used as pathological noise types; the corresponding audio data for each pathological noise type in each nasopharyngeal audio data segment are marked, and the marked audio data is displayed on a screen to assist doctors in assessing the patient's pathological voice. This embodiment uses... This example is used for illustration; no specific limitations are set in this embodiment. It depends on the specific implementation situation.

[0082] To sum up, the pathological voice information processing method in the above embodiment of the present application, according to a plurality of dimension characteristic values under different dimensions, constructs a nasopharynx dimension matrix, and obtains a plurality of dimension-related direction components from the nasopharynx dimension matrix; wherein the dimension-related direction component is used to describe the component information associated on each dimension of the audio information that the doctor mainly focuses on, and the doctor's experience in the clinic is integrated into each dimension of the detected voice information, so that the association between each dimension of the voice information is more intelligent; then the abnormal fluctuation of data on the dimension-related direction component is analyzed in combination with the dimension-related characteristic value, to obtain the phonetic pathological value of the dimension-related direction component; then the value of the abnormal fluctuation on the dimension-related direction component is analyzed, and the phonetic pathological values of different dimension-related direction components are integrated, to obtain the pathological voice degree of the phoneme category; wherein the pathological voice degree is used to describe the obvious degree of pathological characteristics of the patient's pronunciation part when pronouncing the corresponding phoneme category, and can quantify the pathological relationship of the patient when pronouncing different phoneme categories; the accuracy of the patient's voice information recognition is improved, and the early auxiliary evaluation result of the patient's pathological voice can provide a reference value for the doctor. The problem that the accuracy of the analysis result of the pathological voice in the prior art is low, and the early auxiliary evaluation result of the patient's pathological voice can provide a low reference value for the doctor is solved.

[0083] Embodiment two

[0084] Please refer to Figure 2 , which is a pathological voice information processing system proposed in the second embodiment of the present application, the system comprises:

[0085] The nasopharynx sound segment acquisition module 100 is used to acquire a plurality of nasopharynx sound segments of the same phoneme category under different dimensions, and the nasopharynx sound segments correspond to a dimension characteristic value under different dimensions respectively;

[0086] The dimension-related direction component acquisition module 200 is used to construct a nasopharynx dimension matrix by using the nasopharynx sound segments according to a plurality of dimension characteristic values under different dimensions; perform feature analysis on the nasopharynx dimension matrix to obtain a dimension-related feature vector and a dimension-related characteristic value; and obtain a dimension-related direction component in the direction indicated by each dimension-related feature vector from the nasopharynx dimension matrix according to the dimension-related feature vector;

[0087] The pathological voice degree acquisition module 300 is used to analyze the abnormal fluctuation of data on the dimension-related direction component as the pathological performance degree of each dimension-related direction component; combine the dimension-related characteristic value with the pathological performance degree to obtain the phonetic pathological value of each dimension-related direction component; analyze the value of the abnormal fluctuation on the dimension-related direction component, and integrate the phonetic pathological values of different dimension-related direction components to obtain the pathological voice degree of the phoneme category.

[0088] The functions or operation steps realized when the above modules are executed are substantially the same as the method embodiments described above, and will not be described here again.

[0089] Embodiment three

[0090] Another aspect of the present application also provides a readable storage medium, which stores a computer program, and the program realizes the steps of the method described in the above embodiment one when executed by a processor.

[0091] Embodiment four

[0092] Another aspect of the present application also provides an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor realizes the steps of the method described in the above embodiment one when executing the program.

[0093] The technical features of each of the above embodiments can be combined arbitrarily, and in order to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.

[0094] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a list of executable instructions for realizing the logic function, which can be embodied in any computer readable storage medium for use by or in connection with an instruction execution system, apparatus or device, such as a computer-based system, a system including a processor or other system that can fetch instructions from the instruction execution system, apparatus or device and execute the instructions, or in conjunction with these instructions execution system, apparatus or device. For the present specification, the "computer readable storage medium" can be any device that can contain, store, communicate, propagate or transport programs for use by or in connection with an instruction execution system, apparatus or device, or in conjunction with these instructions execution system, apparatus or device.

[0095] More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: an electrical connection having one or more wires (electrical devices), a portable computer diskette (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CD ROM). In addition, the computer readable storage medium can even be paper or other suitable medium on which the program can be printed, because the program can be electronically obtained, for example, by optical scanning of the paper or other medium, followed by editing, interpretation or processing as necessary, and then stored in a computer memory if necessary.

[0096] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the embodiments described above, various steps or methods can be implemented, for example, in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, and in another embodiment, any of the following techniques can be used to implement the hardware used to implement the described functions: discrete logic circuitry having logic gates for implementing logic functions upon data signals, application specific integrated circuits having logic gates for implementing the logic functions on data signals, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.

[0097] In the description of the specification, the description of the terms "one embodiment", "some embodiments", "an example", "a specific example", or "some examples" and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In the specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in one or more embodiments or examples.

[0098] The above-described embodiments only express several implementation manners of the present application, which are described in a more specific and detailed manner, but cannot be understood as a limitation on the patent scope of the present application. It should be noted that, for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present application, which are all within the protection scope of the present application. Therefore, the patent protection scope of the present application should be subject to the appended claims.

Claims

1. A method of processing pathological voice information, characterized by, The method comprises: Obtaining a plurality of nasopharyngeal phonemes of the same phoneme category in different dimensions, wherein the nasopharyngeal phonemes correspond to a dimension feature value in different dimensions respectively; According to the plurality of dimension feature values in different dimensions, constructing a nasopharyngeal dimension matrix by using the nasopharyngeal phonemes; performing feature analysis on the nasopharyngeal dimension matrix to obtain a dimension-related feature vector and a dimension-related feature value; and obtaining a dimension-related direction component in a direction indicated by each dimension-related feature vector from the nasopharyngeal dimension matrix according to the dimension-related feature vector; Analyzing abnormal fluctuations of data on the dimension-related direction component as a pathological performance degree of each dimension-related direction component; taking a value obtained by combining the dimension-related feature value and the pathological performance degree as a phonetic pathological value of each dimension-related direction component; analyzing the value of abnormal fluctuations on the dimension-related direction component, and taking a value obtained by integrating phonetic pathological values of different dimension-related direction components as a pathological voice degree of the phoneme category, comprising: For any one dimension-related direction component, taking a value obtained by combining a dimension feature value on a dimension-related feature vector to which the dimension-related direction component belongs and a pathological performance degree as a phonetic pathological value of the dimension-related direction component; For any one dimension-related direction component, obtaining a component pathological voice degree of the dimension-related direction component according to the pathological performance degree and the phonetic pathological value; Obtaining the component pathological voice degree of each dimension-related direction component; Taking a cumulative mapping result of the component pathological voice degrees of all dimension-related direction components as the pathological voice degree of the phoneme category; The method for obtaining the nasopharyngeal dimension matrix comprises: For any one dimension, constructing a sequence of dimension feature values of the dimension on all nasopharyngeal phonemes, and taking the constructed sequence as a phonetic dimension representation sequence of the dimension; Obtaining the phonetic dimension representation sequence of each dimension; Taking the phonetic dimension representation sequence of each dimension as a column vector, and constructing a matrix, wherein the constructed matrix is taken as the nasopharyngeal dimension matrix; The method for obtaining the dimension-related feature vector and the dimension-related feature value comprises: Performing eigenvalue decomposition on the nasopharyngeal dimension matrix to obtain a plurality of feature vectors and a plurality of feature values, wherein each feature vector is taken as a dimension-related feature vector, and each feature value is taken as a dimension-related feature value; each dimension-related feature vector corresponds to a dimension-related feature value; The method for obtaining the dimension-related direction component comprises: For any one dimension-related feature vector, taking a vector obtained by multiplying the dimension-related feature vector and the nasopharyngeal dimension matrix as a dimension-related direction component in a direction indicated by the dimension-related feature vector; The method for obtaining the pathological performance degree comprises: For any one dimension-related direction component, obtaining a plurality of component abnormal data of the dimension-related direction component; Taking an overall average amplitude of all component abnormal data as the pathological performance degree of the dimension-related direction component.

2. The method of processing pathological voice information according to claim 1, characterized in that, The method for obtaining the component abnormal data comprises: Performing abnormal detection on data on the dimension-related direction component by using an abnormal detection algorithm to obtain a plurality of abnormal data of the dimension-related direction component, wherein each abnormal data is taken as component abnormal data.

3. The method of processing pathologic voice information according to claim 1, wherein, The method for obtaining the component pathological voice degree comprises: The value obtained by combining the pathological manifestation degree of the dimension-related directional component with the speech pathological value is taken as the pathological voice degree of the component.

4. The method of processing pathologic voice information according to claim 1, wherein, The step of analyzing the abnormal fluctuation of the value of the dimension-related directional component and integrating the speech pathological values of different dimension-related directional components to obtain the pathological voice degree of the phoneme category further comprises: Normalizing the pathological voice degree.

5. A system for processing pathological voice information, characterized by The system for implementing the pathological voice information processing method according to any one of claims 1 to 4 comprises: a nasopharyngeal sound segment acquisition module configured to acquire nasopharyngeal sound segments of the same phoneme category under different dimensions, wherein the nasopharyngeal sound segments correspond to a dimension characteristic value under different dimensions respectively; a dimension-related directional component acquisition module configured to construct a nasopharyngeal dimension matrix by using the nasopharyngeal sound segments according to the dimension characteristic values under different dimensions, to acquire dimension-related characteristic vectors and dimension-related characteristic values by performing feature analysis on the nasopharyngeal dimension matrix, and to acquire a dimension-related directional component in the direction indicated by each dimension-related characteristic vector from the nasopharyngeal dimension matrix according to the dimension-related characteristic vector; a pathological voice degree acquisition module configured to analyze abnormal fluctuation of data of the dimension-related directional component to obtain a pathological manifestation degree of each dimension-related directional component, to combine the dimension-related characteristic value with the pathological manifestation degree to obtain a speech pathological value of each dimension-related directional component, to analyze the abnormal fluctuation of the value of the dimension-related directional component, and to integrate the speech pathological values of different dimension-related directional components to obtain the pathological voice degree of the phoneme category.

Citation Information

Patent Citations

  • Examination cheating recognition method and device based on voice recognition and computer equipment

    CN112669820A

  • Health data acquisition and intelligent analysis method

    CN116705337A