Data processing method, respiration sounding capability determination system and medium

By acquiring the energy and fundamental frequency characteristics of speech data and adjusting the vocal energy threshold in conjunction with environmental noise characteristics, the effective vocal range is identified, solving the problem of misjudgment in conventional assessment methods and achieving a more accurate and stable assessment of respiratory vocal capacity.

CN121817809APending Publication Date: 2026-04-10HENAN JIAYU MEDICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HENAN JIAYU MEDICAL TECH CO LTD
Filing Date
2026-02-05
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing speech and cognitive function assessments, conventional assessment methods lead to errors and underestimations of assessment indicators, affecting the accuracy and stability of assessment results. In particular, silent fricatives and continuous counting are easily misjudged as speech interruptions.

Method used

By acquiring the energy and fundamental frequency characteristic data of speech data, the speech energy threshold is dynamically adjusted according to the current environmental noise characteristics. The effective speech range is identified in combination with the task characteristics, and the speech ability index and breathing-speech coordination result are determined based on the effective speech range. A unified data processing flow is adopted to reduce subjectivity.

Benefits of technology

It improves the accuracy and stability of the assessment results, can more accurately identify the effective vocal range, reduce misjudgments, and improve the accuracy and coordination of the assessment of breathing and vocalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121817809A_ABST
    Figure CN121817809A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, a breathing sound production capacity determination system and a medium, and relates to the technical field of sound recognition. And processing the acoustic characteristic data based on the task characteristics of the current breathing sound production function task and the sound production energy threshold to obtain a corresponding effective sound production interval. The sound production energy threshold value is updated according to the real-time environment noise, and the stability and adaptability of effective sound production interval recognition are improved. Considering that fundamental frequency characteristic data of silent friction sound in a normal sound production state does not exist, adaptive processing is realized according to sound production mechanism differences of different function tasks. And determining the sound production capability index and the breathing and sound production coordination result based on the effective sound production interval, determining the corresponding sound production capability index according to different effective sound production intervals, and evaluating the breathing-sound production coordination based on the breathing and sound production coordination result. A unified data processing flow is adopted in the whole process, and compared with conventional manual judgment, subjectivity is avoided, and the accuracy and stability of an evaluation result are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of voice recognition technology, and in particular to a data processing method, a breathing and vocalization ability determination system, and a medium. Background Technology

[0002] In speech and cognitive function assessment, an individual's breathing control and vocalization ability are mainly assessed through tasks such as Maximum Phonation Time (MPT), maximum counting ability, and the duration and ratio of sustained / s / and / z / phonations.

[0003] Two main assessment methods were used. One was manual assessment, where participants manually timed and judged whether breathing occurred or was interrupted. This method was highly subjective, leading to significant differences in results between different assessments. The other method used a uniform fixed energy threshold and fundamental frequency as necessary conditions for effective vocalization. However, considering that the fundamental frequency signal for the silent fricative / s / is not present during vocalization, it was incorrectly judged as invalid vocalization. Furthermore, during / z / or continuous counting, although the subject was vocalizing, there were short-term drops in fundamental frequency and energy fluctuations, which were incorrectly judged as vocal interruptions. Both of these assessment methods resulted in the underestimation of assessment indicators, affecting the accuracy and stability of the assessment results.

[0004] Therefore, improving the accuracy and stability of evaluation results is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] The purpose of this application is to provide a data processing method, a breathing and vocalization ability determination system and medium to solve the problem of underestimation of evaluation indicators under conventional evaluation methods, which reduces the accuracy and stability of evaluation results.

[0006] To address the aforementioned technical problems, this application provides a data processing method, comprising: Acquire acoustic feature data and the corresponding current breathing and vocalization function task obtained by feature extraction processing of the speech data to be tested; wherein, the acoustic feature data includes at least energy feature data and fundamental frequency feature data; Based on the task characteristics and vocal energy threshold of the current breathing and vocalization function task, the acoustic feature data is processed to obtain the corresponding effective vocalization range; wherein, the vocal energy threshold is pre-processed based on the current environmental noise energy feature data of the speech data. The vocal capacity index and the results of breath-voice coordination are determined based on the effective vocal range.

[0007] On the one hand, when the current breathing and vocalization task is a silent friction phonation task, the acoustic feature data is processed based on the task characteristics and vocal energy threshold of the current breathing and vocalization task to obtain the corresponding effective vocalization range, including: The data type of acoustic feature data is determined based on the task characteristics of the silent friction sound generation function task; wherein, the data type includes the data type of the energy feature data; Filter the energy feature data that exceeds the vocal energy threshold to select a first target energy feature data. If the time corresponding to each first target energy feature data is continuous and the continuous time is greater than the first preset continuous time, the sound emission interval corresponding to each first target energy feature data is taken as the effective sound emission interval.

[0008] On the other hand, when the current breathing and vocalization function task is a continuous vocalization function task, the acoustic feature data is processed based on the task characteristics and vocal energy threshold of the current breathing and vocalization function task to obtain the corresponding effective vocalization range, including: The data type of acoustic feature data is determined based on the task characteristics of the continuous sound emission function task; wherein, the data type includes the data type of energy feature data and the data type of fundamental frequency feature data; The corresponding preset range of fundamental frequency is determined in advance based on the subject's age parameter corresponding to the voice data; Filter the energy feature data that exceeds the vocal energy threshold to select a first target energy feature data. The corresponding first target fundamental frequency feature data is determined based on the target speech data corresponding to the first target energy feature data; Select second target base frequency feature data that fall within the corresponding preset range of base frequency from each first target base frequency feature data; When the time corresponding to each second target fundamental frequency characteristic data is continuous and the continuous time is greater than the second preset continuous time, the sound range corresponding to each second target fundamental frequency characteristic data is taken as the effective sound range.

[0009] On the other hand, when the current breathing and vocalization function task is a continuous counting function task, the acoustic feature data is processed based on the task characteristics and vocal energy threshold of the current breathing and vocalization function task to obtain the corresponding effective vocalization range, including: The data type of acoustic feature data is determined based on the task characteristics of the continuous counting function; wherein, the data type includes the data type of energy feature data and the data type of fundamental frequency feature data; Filter the energy feature data that exceeds the vocal energy threshold to select a first target energy feature data. When the time corresponding to the first target energy feature data is continuous and the continuous time is greater than the third preset continuous time, the speech data corresponding to the first target energy feature data is aggregated to obtain a speech segment. Each sound segment is labeled as a discrete sound event, and the time position corresponding to each discrete sound event is marked to obtain the event label; The event markers are sorted to obtain a sequence of vocal events; If the time interval between adjacent sound events in the sound event sequence is less than a preset time interval, then the counting process corresponding to the sound event sequence is determined to be the initial counting process. The corresponding preset range of fundamental frequency is determined based on the subject's age parameter corresponding to the voice data of the initial counting process; Select third target base frequency feature data that falls within the corresponding preset range of base frequency from the base frequency feature data during the initial counting process; When the time corresponding to each third target fundamental frequency characteristic data is continuous and the continuous time is greater than the third preset continuous time, the sound interval corresponding to each third target fundamental frequency characteristic data is taken as the effective sound interval.

[0010] On the other hand, vocal ability indicators are determined based on the effective vocal range, including: Mark the time axis corresponding to the sound frames within the effective sound range to determine the continuous effective sound range; The maximum effective vocal range is obtained by extracting the continuous effective vocal range based on the start and end times and duration of the vocalization. Within the maximum effective vocal range, the actual vocal ability index is obtained by processing the speech data within the maximum effective vocal range according to the preset vocal ability indexes under the breathing vocal function task.

[0011] On the other hand, the maximum effective vocal range is obtained by extracting the continuous effective vocal range based on the start and end times and duration of the vocalization, including: Within the continuous effective sound emission interval, if the corresponding actual start and end time is greater than or equal to the fourth preset continuous time, and the number of events under the corresponding duration is greater than or equal to the preset number of events, then the continuous effective sound emission interval to which the corresponding actual start and end time is greater than or equal to the fourth preset continuous time, and the number of events under the corresponding duration is greater than or equal to the preset number of events, shall be taken as the maximum effective sound emission interval.

[0012] On the other hand, the process of determining the result of the breath-voice coordination includes: Determine the target duration of the silent zone within the maximum effective vocal range; If the duration of the target exceeds the fifth preset continuous time, it is determined that there is a pause in speech or a gap in speech within the maximum effective vocal range, and the result of the breathing-voice coordination is insufficient. If no silent interval is detected, or if the target duration of the silent interval is less than or equal to the fifth preset continuous time, then the breathing and vocalization coordination result is determined to be good.

[0013] On the other hand, the process of determining the sound energy threshold includes: Obtain the current ambient noise data of the location where the voice data is located; Multiple current environmental noise energy feature data are obtained by feature extraction based on the current environmental noise data; The average environmental noise energy characteristic data is obtained by averaging multiple current environmental noise energy characteristic data. The sound energy threshold is determined based on the average environmental noise energy characteristic data. Correspondingly, determining the sound energy threshold based on the average environmental noise energy characteristic data includes: Obtain a first empirical coefficient and a second empirical coefficient; sum the average environmental noise energy characteristic data and the first empirical coefficient to obtain the sound energy threshold; Alternatively, obtain a second empirical coefficient; process multiple current environmental noise energy characteristic data by standard deviation to obtain standard environmental noise energy characteristic data; obtain first environmental noise energy characteristic data based on the second empirical coefficient and the standard environmental noise energy characteristic data; and sum the average environmental noise energy characteristic data and the first environmental noise energy characteristic data to obtain the sound energy threshold.

[0014] To address the aforementioned technical problems, this application also provides a system for determining respiratory vocalization ability, comprising: Memory, used to store computer programs; A processor for implementing the steps of the data processing method as described above when executing the computer program.

[0015] To address the aforementioned technical problems, this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the data processing method described above.

[0016] This application provides a data processing method. First, it acquires acoustic feature data obtained by feature extraction processing of the speech data to be tested, as well as the current respiratory phonation function task corresponding to the speech data. The feature extraction process extracts at least energy feature data and fundamental frequency feature data. Energy feature data is used for phonation detection and continuity analysis, while fundamental frequency feature data is used for determining phonation activity; however, fundamental frequency feature data is allowed to be missing in the absence of sound or weak sound. Regarding the respiratory phonation function task, it is considered that silent friction sounds are easily misjudged as invalid phonation in normal phonation states because the fundamental frequency itself is absent; and that in voiced sounds or continuous counting, phonation interruption may be misjudged due to respiratory fluctuations and transient vocal cord instability. This application determines the corresponding respiratory phonation function task for the current speech data to facilitate separate evaluation of the specific respiratory phonation function task. Second, based on the task characteristics and phonation energy threshold of the current respiratory phonation function task, the acoustic feature data is processed to obtain the corresponding effective phonation range. The vocal energy threshold here is obtained based on the energy characteristic data of the current ambient noise. Compared with conventional solutions that use the same fixed energy threshold for different ambient noise levels, leading to invalid judgments, this application updates the vocal energy threshold according to real-time ambient noise, improving the stability and adaptability of effective vocal range identification. Furthermore, based on the task characteristics of the current breathing vocalization function task, and considering the absence of fundamental frequency characteristic data for silent frictional sound under normal vocalization conditions, the processing method for the determined effective vocal range differs. This allows for adaptive processing based on the differences in vocalization mechanisms of different functional tasks, enabling the acquisition of corresponding effective vocal ranges through dedicated processing methods for different vocalization tasks. While conventional solutions also extract energy and fundamental frequency characteristic data, this application adds its own processing methods to the vocal energy threshold based on the task characteristics of different breathing vocalization function tasks, achieving differentiated processing. Finally, vocal ability indicators and breathing-vocalization coordination results are determined based on the effective vocal range. This allows for the determination of corresponding vocal ability indicators based on different effective vocal ranges, improving accuracy, while also evaluating breathing-vocalization coordination based on the determined breathing-vocalization coordination results. The entire process employs a unified data processing workflow, which, compared to conventional human judgment, improves the accuracy and stability of the evaluation results by avoiding subjectivity.

[0017] In addition, this application also provides a breathing and vocalization ability determination system and medium, which have the same beneficial effects as the data processing method described above. Attached Figure Description

[0018] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 A flowchart illustrating a data processing method provided in an embodiment of this application; Figure 2 A flowchart illustrating another data processing method provided in the embodiments of this application; Figure 3 A structural diagram of a data processing apparatus provided in an embodiment of this application; Figure 4 This is a structural diagram of a breathing and vocalization ability determination system provided in an embodiment of this application. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0021] The core of this application is to provide a data processing method, a breathing and vocalization ability determination system and medium to solve the problem of underestimation of evaluation indicators under conventional evaluation methods, which reduces the accuracy and stability of evaluation results.

[0022] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] Conventional speech and cognitive ability assessments employ two methods: one involves professionals manually timing and judging whether breathing or interruptions occur; the other automatically detects the start and end points of speech based on energy thresholds, estimates duration, and determines phonation based on the presence of frequencies, using a uniform phonation judgment logic for / s / and / z / . The manual assessment method is highly subjective, with significant differences in results among assessors. It lacks the ability to quantify fine-grained features such as loudness, fundamental frequency, and rhythm, and is difficult to automate and scale on ordinary terminal devices. Conventional automated speech assessment uses a uniform phonation judgment logic, relying on a fixed energy threshold and fundamental frequency presence as necessary conditions for effective phonation, without differentiating the acoustic characteristics of different phonation types. For example, the silent fricative / s / does not produce a stable fundamental frequency under normal phonation conditions; its acoustic characteristics are mainly reflected in continuous broadband noise energy. In contrast, the audible fricative / z / , vowels, and counting tasks rely on stable vocal cord vibration, exhibiting a continuous fundamental frequency structure.

[0024] On the one hand, during the / s / vocalization process, the fundamental frequency is not present, making it easy to misjudge as invalid vocalization. On the other hand, during / z / or continuous counting, even if the subject is still vocalizing normally, short-term drops in fundamental frequency or energy fluctuations may occur due to breathing fluctuations, momentary instability of the vocal cords, or algorithm estimation errors, which may be incorrectly judged as vocal interruptions. This brief fluctuation is often considered the end of vocalization, directly truncating the longest continuous vocalization interval, leading to an underestimation of the longest duration, s / z ratio, and related ability indicators, affecting the accuracy and stability of the evaluation results. The data processing method provided in this application can solve the above-mentioned technical problems.

[0025] Figure 1 A flowchart of a data processing method provided in an embodiment of this application is shown below. Figure 1 As shown, the method includes: S11: Obtain acoustic feature data and the corresponding current breathing and vocalization function task obtained by feature extraction processing of the speech data to be tested; Among them, acoustic feature data includes at least energy feature data and fundamental frequency feature data; S12: Based on the task characteristics and vocal energy threshold of the current breathing and vocalization function task, the acoustic feature data is processed to obtain the corresponding effective vocalization range; Among them, the sound energy threshold is obtained in advance based on the environmental noise energy characteristic data of the current environment in which the speech data is located; S13: Determine the vocal capacity index and the results of breath-voice coordination based on the effective vocal range.

[0026] Specifically, in step S11, acoustic feature data is obtained by extracting features from the speech data to be tested. This preprocessing operation can be performed on either single-channel or multi-channel speech data from the subject; no limitation is made here. This embodiment considers that vocal ability indicators, such as longest vocal duration, maximum counting ability, and the duration and ratio of continuous / s / and / z / pronunciations, are all based on the temporal continuity acoustic feature analysis of a single subject, without involving sound source localization, spatial information, or multi-channel fusion processing. Therefore, single-channel speech data is sufficient and more consistent with common microphone or terminal acquisition methods in practical application scenarios. If multi-channel speech data is used, any one channel can be selected as the analysis object. After acquiring the single-channel speech signal, preprocessing operations are performed. These preprocessing operations can include DC component removal, amplitude normalization, and optional filtering, etc., which are not limited here and can be set according to the actual situation.

[0027] Subsequently, the preprocessed speech data is segmented into frames according to preset frame length and frame shift parameters to obtain temporal frame-level speech data. Feature extraction is then performed on the frame-level speech data. This feature extraction method can be the same as or different from conventional methods. However, it should be noted that the extracted feature data includes acoustic feature data, at least energy feature data and fundamental frequency feature data. Energy feature data is calculated from the waveform corresponding to the speech data, reflecting quantitative values ​​of sound intensity, vibration amplitude, and airflow magnitude, characterizing the loudness and energy distribution of the sound, and used for sound production detection and continuity analysis. Fundamental frequency feature data is extracted from the speech data, reflecting a set of quantitative values ​​reflecting the speed of vocal cord vibration, representing the fundamental frequency of the periodic opening and closing vibration of the vocal cords during phonation. This can be achieved using autocorrelation methods or the YIN fundamental tone detection algorithm, etc., without limitation. Fundamental frequency feature data is allowed to be missing in the absence of sound or in weak sound conditions.

[0028] When using the autocorrelation method for extraction, the preprocessed speech signal is segmented and windowed, and the short-time autocorrelation function is calculated frame by frame: the current frame signal is correlated with its shifted signal with different delay times to obtain the autocorrelation curve. In the autocorrelation curve, the number of delay points corresponding to the first maximum peak is the period of vocal cord vibration. Dividing the sampling rate by this period yields the fundamental frequency F0 of the current frame. Signals without obvious periodicity, such as unvoiced or noisy frames, are determined to have no effective fundamental frequency. Through frame-by-frame calculation, the fundamental frequency sequence, mean fundamental frequency, fundamental frequency range, and fundamental frequency continuity of the entire speech signal are obtained.

[0029] When using the YIN pitch detection algorithm for extraction, a difference function and cumulative mean normalization are introduced on top of autocorrelation. First, the differential autocorrelation function of the speech frame is calculated, and then cumulative mean normalization is used to suppress noise and harmonic interference, improving the stability and accuracy of pitch detection. In the normalized function curve, the minimum delay point that meets the threshold condition is selected as the pitch period, and the current frame's pitch F0 is calculated. Frames with silent segments, unvoiced segments, or no obvious periodicity are marked as invalid. The final output includes the pitch sequence, pitch presence, pitch fluctuation, and pitch range of consecutive frames.

[0030] Energy feature data can be frame energy feature data or root mean square energy feature data. Extracting these feature data does not involve mixing or fusing different acoustic features, but rather calculating multiple fundamental acoustic features in parallel and independently on the same time frame. Energy features and fundamental frequency features are calculated separately for each frame of the speech signal. Each feature data point is numerically independent and stored separately. However, in subsequent processing, different features are selectively used based on the test task type and judgment rules. For example, only energy features are used in a silent fricative sound production task, while both energy and fundamental frequency features are used in a spoken sound production task. Therefore, the extraction in this step only represents parallel computation of multiple features, not mixing or weighting features.

[0031] The speech data corresponds to the current respiratory and vocal function task. This task simultaneously examines respiratory support, vocal cord vibration, and articulation coordination. Like the MPT (Master of Prostate Test), it belongs to the respiratory and vocal function test. It mainly includes tasks related to silent fricative phonation, sustained phonation, and continuous counting, and may also include other tasks, which are not limited here. The silent fricative phonation and sustained phonation tasks are static, steady-state phonation maintenance tests, while the continuous counting task can be categorized as an extended dynamic continuous phonation task.

[0032] Regarding the task of extracting speech data corresponding to the current breathing and vocalization function, pre-test calibration can be performed to clarify the test instructions and target sound / content. Specifically, it can also be distinguished using fundamental frequency characteristic data, or spectral characteristics, time domain characteristics, and energy characteristics. There are no restrictions here; it can be set according to the actual situation.

[0033] In step S12, the acoustic feature data is processed based on the task characteristics and vocal energy threshold of the current breathing and vocalization task to obtain the corresponding effective vocalization range. Due to the differences in each breathing and vocalization task, for example, the absence of a fundamental frequency feature is not required for the silent friction phonation task, so only the energy feature data and vocal energy threshold are compared. For the continuous vocalization task, after comparing the vocal energy threshold, the fundamental frequency feature data is also used for determination, maintaining relative continuity over time to ensure vocal stability. For the continuous counting task, in addition to the energy feature data and fundamental frequency feature data, the continuity of the counting process also needs to be determined. Based on the differences between each task, corresponding processing methods are used to obtain the respective effective vocalization ranges.

[0034] The effective vocal range refers to a segment of time within which a speech signal is in a true vocal state. It is used to exclude silence, environmental noise, or invalid vocalizations, retaining only the vocalization process that can be used for ability assessment. It can exist in both silent and vocalized speech signals, and is therefore distinguished between silent and vocalized breathing function tasks.

[0035] Regarding the vocal energy threshold, it is adaptively calculated based on the current environmental noise energy characteristic data. Specifically, before formal functional testing, a recording is made for a period of time, and the audio is analyzed. The threshold is obtained by statistically analyzing the energy values ​​of multiple consecutive frames of speech signals. If the current energy characteristic data exceeds the vocal energy threshold, it is judged as a valid vocal range; otherwise, it is considered a silent or noise frame, thereby improving the system's adaptability to different recording environments. Here, "silent" refers to a situation where there is no sound at all. The current environmental noise energy characteristic data is collected in real time from actual environmental noise levels. The corresponding vocal energy threshold varies depending on the time of day (morning, noon, evening), which improves the accuracy of judgment compared to the conventional approach that uses a fixed vocal energy threshold.

[0036] Step S13 involves determining vocal ability indicators and breath-voice coordination results based on the effective vocal range. Based on the effective vocal range, the effective vocal range with the longest duration needs to be extracted to ensure the accuracy of subsequent determination of vocal ability indicators while also improving stability.

[0037] Vocal performance indicators are determined within the longest effective vocal range, such as the longest vocal duration, which characterizes the subject's respiratory support and sustained vocal capacity; the ratio of the longest duration of silent friction rub to the longest duration of audible friction rub, which reflects the state of vocal function; and the longest continuous effective vocal duration in a counting task, which can further analyze the stability of the counting rhythm and vocal interruptions. Other vocal performance indicators can also be used, without limitation, and can be set according to the actual situation.

[0038] The results of breath-vocal coordination directly affect the vocal performance of subjects during continuous phonation or counting tasks, as respiratory support capacity and breath-vocal coordination directly impact their vocal performance. Under normal circumstances, only brief pauses caused by speech structure occur during continuous phonation or counting, without prolonged silent interruptions. By detecting the duration of continuous silent intervals in the speech signal, abnormal silent segments exceeding the normal range of speech pauses are identified. When the duration of a silent interval exceeds a preset threshold, it is determined that there is a significant interruption in ventilation or respiratory support, thus achieving an objective assessment of breath-vocal coordination. This method can effectively distinguish between normal rhythmic pauses in speech and abnormal respiratory interruptions, improving the reliability of breath coordination assessment results.

[0039] This application provides a data processing method. First, it acquires acoustic feature data obtained by feature extraction processing of the speech data to be tested, as well as the current respiratory vocalization function task corresponding to the speech data. The feature extraction process extracts at least energy feature data and fundamental frequency feature data. Energy feature data is used for vocalization detection and continuity analysis, while fundamental frequency feature data is used for determining audible vocalization; however, fundamental frequency feature data is allowed to be missing in the absence of sound or weak sound. Regarding the respiratory vocalization function task, it is considered that silent frictional sounds are easily misjudged as invalid vocalization in normal vocalization because the fundamental frequency itself is absent; and that in voiced sounds or continuous counting, vocalization may be misjudged as interrupted due to respiratory fluctuations and transient vocal cord instability. This application determines the corresponding respiratory vocalization function task for the current speech data to facilitate separate evaluation of the specific respiratory vocalization function task. Secondly, based on the task characteristics and vocalization energy threshold of the current respiratory vocalization function task, the acoustic feature data is processed to obtain the corresponding effective vocalization range. The vocal energy threshold here is obtained based on the energy characteristic data of the current ambient noise. Compared with conventional solutions that use the same fixed energy threshold for different ambient noise levels, leading to invalid judgments, this application updates the vocal energy threshold according to real-time ambient noise, improving the stability and adaptability of effective vocal range identification. Furthermore, based on the task characteristics of the current breathing vocalization function task, and considering the absence of fundamental frequency characteristic data for silent frictional sound under normal vocalization conditions, the processing method for the determined effective vocal range differs. This allows for adaptive processing based on the differences in vocalization mechanisms of different functional tasks, enabling the acquisition of corresponding effective vocal ranges through dedicated processing methods for different vocalization tasks. While conventional solutions also extract energy and fundamental frequency characteristic data, this application adds its own processing methods to the vocal energy threshold based on the task characteristics of different breathing vocalization function tasks, achieving differentiated processing. Finally, vocal ability indicators and breathing-vocalization coordination results are determined based on the effective vocal range. This allows for the determination of corresponding vocal ability indicators based on different effective vocal ranges, improving accuracy, while also evaluating breathing-vocalization coordination based on the determined breathing-vocalization coordination results. The entire process employs a unified data processing workflow, which, compared to conventional human judgment, improves the accuracy and stability of the evaluation results by avoiding subjectivity.

[0040] In some embodiments, when the current respiratory vocalization task is a silent friction phonation task, acoustic feature data is processed based on the task characteristics and vocal energy threshold of the current respiratory vocalization task to obtain the corresponding effective vocal range, including: The data type of acoustic feature data is determined based on the task characteristics of the silent friction sound generation function task; among which, the data type includes the data type of energy feature data; Filter the energy feature data that exceeds the vocal energy threshold to identify the first target energy feature data. If the time corresponding to each first target energy characteristic data is continuous and the continuous time is greater than the first preset continuous time, the sound emission interval corresponding to each first target energy characteristic data is taken as the effective sound emission interval.

[0041] Specifically, the data type of acoustic feature data is determined based on the task characteristics of the silent friction sound generation function. The purpose is to determine which feature data to use for subsequently judging the valid sound generation interval; here, only energy feature data is considered. From the energy feature data, the first target energy feature data corresponding to each exceeding the sound generation energy threshold are selected. For accuracy, the sound generation interval corresponding to each first target energy feature data needs to be continuous, and if the continuous sound generation time is greater than a first preset continuous time, then the corresponding sound generation interval is considered the valid sound generation interval.

[0042] Since the fundamental frequency feature can be missing in the silent or weak sound generation task, judging the fundamental frequency feature may result in misjudgment and increase the judgment time.

[0043] The process for determining the effective sound range during the silent fricative sound generation task provided in this embodiment uses the energy generation threshold and time continuity as the basis for judgment. It can accurately adapt to the acoustic characteristics of unvoiced sound generation, eliminate invalid segments such as noise, discontinuity, and air leakage, and avoid misjudgment and omission due to the lack of fundamental frequency. This improves the accuracy, robustness and applicability of the effective sound generation determination under the silent fricative sound generation task, and more realistically reflects the subject's airflow control and articulation ability.

[0044] In some embodiments, when the current respiratory phonation function task is a sustained vocalization function task, acoustic feature data is processed based on the task characteristics and vocal energy threshold of the current respiratory phonation function task to obtain the corresponding effective vocalization range, including: The data type of acoustic feature data is determined based on the task characteristics of the continuous sound emission function task; among which, the data type includes the data type of energy feature data and the data type of fundamental frequency feature data; The corresponding fundamental frequency preset range is determined in advance based on the subject's age parameter corresponding to the voice data; Filter the energy feature data that exceeds the vocal energy threshold to identify the first target energy feature data. The corresponding first target fundamental frequency feature data is determined based on the target speech data corresponding to the first target energy feature data; Select second target base frequency feature data that fall within the corresponding preset range of base frequency from each first target base frequency feature data; When the time corresponding to each second target fundamental frequency characteristic data is continuous and the continuous time is greater than the second preset continuous time, the sound range corresponding to each second target fundamental frequency characteristic data is taken as the effective sound range.

[0045] Specifically, based on the task characteristics of sustained vocalization, the data type of acoustic features is determined, including energy feature data and audio feature data. Fundamental frequency feature data exhibits significant and regular physiological changes with age: the fundamental frequency is generally high in childhood, decreases sharply in males after puberty, remains relatively stable in adulthood, and then increases or decreases overall in old age due to vocal cord aging and changes in muscle tone. The normal physiological fundamental frequency range varies significantly among subjects of different ages, and there is no uniformly applicable fixed fundamental frequency range.

[0046] The process for determining the first target energy feature data is the same as in the above embodiment, and will not be repeated here. Then, based on the target speech data corresponding to the first target energy feature data, the corresponding first target fundamental frequency feature data is determined. Here, the second target fundamental frequency feature data within the corresponding preset fundamental frequency range is further filtered out. To ensure accuracy, the phonation time corresponding to each second target fundamental frequency feature data needs to be continuous, and if the continuous phonation time is greater than the second preset continuous time, the corresponding phonation interval is taken as the valid phonation interval.

[0047] The determination process for the effective sound range during continuous sound emission tasks provided in this embodiment, by adding a preset range of fundamental frequency and time continuity constraints on the basis of sound energy threshold determination, can ensure the stability of pitch, vibration pattern and sound continuity of continuous sound emission, eliminate invalid segments such as abnormal fundamental frequency, discontinuity, and instability, ensure that the effective sound segments meet the steady-state sound emission requirements, and improve the accuracy of sound emission stability determination and acoustic feature analysis.

[0048] In some embodiments, when the current breathing and vocalization task is a continuous counting task, the acoustic feature data is processed based on the task characteristics and vocal energy threshold of the current breathing and vocalization task to obtain the corresponding effective vocalization range, including: The data type of acoustic feature data is determined based on the task characteristics of the continuous counting function; among which, the data type includes the data type of energy feature data and the data type of fundamental frequency feature data; Filter the energy feature data that exceeds the vocal energy threshold to identify the first target energy feature data. If the time corresponding to the first target energy feature data is continuous and the continuous time is greater than the third preset continuous time, the speech data corresponding to the first target energy feature data is aggregated to obtain a speech segment. Each sound segment is labeled as a discrete sound event, and the time position corresponding to each discrete sound event is marked to obtain the event label; The event tags are sorted to obtain a sequence of vocal events; If the time interval between adjacent sound events in a sound event sequence is less than a preset time interval, then the counting process corresponding to the sound event sequence is determined to be the initial counting process. The corresponding preset range of fundamental frequency is determined based on the subject's age parameter corresponding to the voice data of the initial counting process; Select the third target fundamental frequency characteristic data that falls within the corresponding preset range of fundamental frequency from the fundamental frequency characteristic data during the initial counting process; If the time corresponding to each third target fundamental frequency characteristic data is continuous and the continuous time is greater than the third preset continuous time, the sound range corresponding to each third target fundamental frequency characteristic data is taken as the effective sound range.

[0049] Specifically, for the continuous counting task, it is not required that the speech signal maintain strict frame-level continuity at the acoustic level. Instead, based on the monitoring results of discrete sound events during the counting process, and under the premise of allowing brief normal pauses, it is determined whether adjacent sound events continue to occur within a reasonable time range, thereby achieving effective judgment of the counting task.

[0050] First, discrete vocal events are detected primarily based on energy feature data. Based on frame-level energy feature data, consecutive vocal frames exceeding a vocal energy threshold are aggregated. That is, in the same manner as the previous embodiment, after filtering out the first target energy feature data, if the corresponding time is continuous and the continuous time is greater than a third preset continuous time, the speech data corresponding to the first target energy feature data is aggregated to obtain vocal segments. Each vocal segment corresponds to a discrete vocal event, and its time position is extracted as an event marker, ultimately resulting in a sequence of vocal events arranged in chronological order.

[0051] Within a sequence of vocal events, if the time interval between adjacent vocal events is less than a preset time interval, the counting process corresponding to the vocal event sequence is determined to be continuous, i.e., the initial counting process. Allowing for normal speech pauses, the preset fundamental frequency range corresponding to the fundamental frequency feature data is used as an auxiliary constraint. This is only used in event-level judgment to verify whether the vocal events have reasonable fundamental frequency characteristics, to exclude abnormal noise or non-speech events, and does not require the speech signal to maintain strict fundamental frequency continuity at the frame level. Specifically, the third target fundamental frequency feature data within the corresponding preset fundamental frequency range is selected. If the corresponding time is continuous and the continuous time is greater than the third preset continuous time, the vocal intervals corresponding to each of these third target fundamental frequency feature data are considered valid vocal intervals.

[0052] The determination process for valid vocal intervals in the continuous counting task provided in this embodiment detects discrete vocal events primarily based on energy characteristics. Under the premise of allowing normal physiological pauses, it analyzes the time interval between adjacent vocal events to determine continuity by combining the existence of the fundamental frequency and reasonable range constraints. It does not require strict fundamental frequency continuity at the frame level, which can fit the actual characteristics of natural vocalization in continuous speech. It effectively eliminates discontinuities such as excessively long pauses, interruptions, repetitions, and omissions, and avoids being misjudged as invalid due to short-term interruptions of the fundamental frequency caused by normal inter-word breathing, short pauses, and articulation switching. This improves the rationality, robustness, and accuracy of vocal continuity determination in the continuous counting task, and more realistically reflects the subject's breathing-vocalization-articulation coordination ability in continuous speech.

[0053] In some embodiments, determining a vocal capacity index based on an effective vocal range includes: Mark the timeline corresponding to the sound frames within the effective sound range to determine the continuous effective sound range; The maximum effective vocal range is obtained by extracting the continuous effective vocal range based on the start and end times and duration of the vocalization. Within the maximum effective vocal range, the actual vocal ability index is obtained by processing the speech data within the maximum effective vocal range according to the preset vocal ability indexes under the breathing vocal function task.

[0054] Specifically, valid and invalid vocal frames are marked, and the marking results are analyzed on the timeline to identify continuous valid vocal intervals, obtaining their start and end times and durations. The longest-lasting valid vocal interval, i.e., the maximum valid vocal interval, is then extracted. Within the identified maximum valid vocal interval, the speech data is processed according to the preset vocal ability indicators under the breathing vocal function task to obtain the actual vocal ability indicators.

[0055] Vocal performance indicators include, but are not limited to: longest vocal duration, used to characterize the subject's respiratory support and sustained vocal capacity; the ratio of the longest vocal duration of silent friction rub to that of audible friction rub, used to reflect the state of vocal function; and the longest continuous effective vocal duration in a continuous counting task, which can further analyze the stability of the counting rhythm and the occurrence of vocal interruptions.

[0056] The duration, event interval, and statistical characteristics of different test tasks are calculated to output corresponding vocal ability evaluation indicators. For example, when the test task is to measure the longest vocal time, the maximum effective vocal range is used as the target range. The duration of continuous vocalization is calculated by the difference between the start and end times of this range. This duration is the maximum vocal time, which is used to characterize the subject's respiratory support ability and sustained vocalization ability.

[0057] If the current test task is the s / z ratio, the above implementation method is executed for the / s / vocalization task and the / z / vocalization task respectively. Then, for each vocalization type, the longest duration corresponding to the longest continuous effective vocalization interval is calculated independently. Finally, the ratio of the longest duration of / s / to the longest duration of / z / is calculated to obtain the s / z ratio.

[0058] When the test task is to measure the maximum counting ability, the longest event sequence belonging to the same continuous counting process is first identified based on the detected discrete event sequence and the determination result of the maximum effective vocal range. The longest duration of the continuous counting process is accelerated by the difference between the start time and the end time of the event sequence, which is used as the maximum counting ability index.

[0059] The process of determining the vocal ability indicators provided in this embodiment can rely on high-quality and effective speech data to achieve a quantitative representation of vocal ability, improve the accuracy and reliability of indicators such as longest voice duration, voice duration ratio, and continuous voice duration, comprehensively reflect the subject's breathing support, airflow control and continuous speech coordination function, and make the evaluation results more objective.

[0060] In some embodiments, the maximum effective vocal range is obtained by extracting the continuous effective vocal range based on the start and end times and duration of the vocalization, including: Within a continuous effective sound interval, if the corresponding actual start and end time is greater than or equal to the fourth preset continuous time, and the number of events under the corresponding duration is greater than or equal to the preset number of events, then the continuous effective sound interval to which the corresponding actual start and end time is greater than or equal to the fourth preset continuous time, and the number of events under the corresponding duration is greater than or equal to the preset number of events, shall be taken as the maximum effective sound interval.

[0061] Specifically, if within a continuous valid sound interval, the actual start and end time is greater than or equal to the fourth preset continuous time, and the corresponding number of events is greater than or equal to the preset number of events, then this portion of the continuous valid sound interval is considered the maximum valid sound interval. If only a continuous valid sound interval less than the fourth preset continuous time or with a corresponding number of events less than the preset number of events is detected, the test can be deemed invalid, and no capability assessment result can be output.

[0062] The process for determining the maximum effective vocal range provided in this embodiment can accurately identify continuous effective vocal ranges, determine their start and end times and duration, and focus on extracting the effective vocal range with the longest duration, ensuring that only the most stable and representative vocal data are used for subsequent evaluation.

[0063] In some embodiments, the process of determining the result of breath-voice coordination includes: Determine the target duration of the silent zone within the maximum effective vocal range; If the target duration exceeds the fifth preset continuous time, it is determined that there is a pause in speech or a gap in speech within the maximum effective vocal range, and that the breathing and vocal coordination result is insufficient. If no silent interval is detected or the target duration of the silent interval is less than or equal to the fifth preset continuous time, then the breathing and vocalization coordination result is determined to be good.

[0064] Specifically, within the longest continuous effective vocal interval, the speech data is marked with frame-level validity, and the target duration of continuous silent intervals is detected on the time axis. When a silent interval with a target duration exceeding the fifth preset continuous time is detected, the silent interruption is considered to exceed the physiological range of normal speech pauses or vocal gaps, thus indicating that the subject may have deficiencies in respiratory support or respiratory-vocal coordination; if no such abnormal silent intervals are detected, respiratory-vocal coordination is considered to be good.

[0065] This analysis focuses on whether there are abnormal interruptions in vocalization, rather than a simple comparison between different time periods, thus providing a more objective reflection of the subject's breathing-vocal coordination ability in continuous vocalization or counting tasks.

[0066] The process for determining the respiratory-vocal coordination result provided in this embodiment detects the duration of continuous silent intervals in speech data and identifies abnormal silent segments that exceed the normal range of speech pauses. When the duration of a silent interval exceeds a preset threshold, it is determined that there is a significant interruption in ventilation or respiratory support, thereby achieving an objective assessment of respiratory-vocal coordination. This effectively distinguishes between normal rhythmic pauses in speech and abnormal respiratory interruptions, improving the reliability of the respiratory coordination assessment results.

[0067] In some embodiments, the process of determining the acoustic energy threshold includes: Obtain the current ambient noise data of the speech data; Multiple current environmental noise energy feature data are obtained by extracting features from the current environmental noise data; The average environmental noise energy characteristic data is obtained by averaging multiple current environmental noise energy characteristic data. The sound energy threshold is determined based on the average environmental noise energy characteristic data. Correspondingly, the sound energy threshold is determined based on the average environmental noise energy characteristic data, including: Obtain the first empirical coefficient and the second empirical coefficient; sum the average environmental noise energy characteristic data and the first empirical coefficient to obtain the sound energy threshold; Alternatively, obtain a second empirical coefficient; process multiple current environmental noise energy characteristic data by standard deviation to obtain standard environmental noise energy characteristic data; obtain first environmental noise energy characteristic data based on the second empirical coefficient and standard environmental noise energy characteristic data; and sum the average environmental noise energy characteristic data and the first environmental noise energy characteristic data to obtain the sound energy threshold.

[0068] Specifically, the calculation formula is: T_energy=μ_noise+K; or, T_energy=μ_noise+α·σ_noise; where T_energy is the sound energy threshold, μ_noise is the average environmental noise energy characteristic data, K is the first empirical coefficient, α is the second empirical coefficient, σ_noise is the standard environmental noise energy characteristic data, and α·σ_noise is the first environmental noise energy characteristic data.

[0069] The sound energy threshold determination process provided in this embodiment distinguishes between effective sound emission states and silent or noisy states, thereby improving the system's adaptability to different recording environments. It reduces the impact of different devices and environmental noise on sound emission detection results, improving the algorithm's robustness.

[0070] In the entire data processing process described above, after determining the valid sound frames, each frame is marked as either a "valid sound frame" or an "invalid frame," thus forming a binary time series arranged in chronological order. By traversing this time series, the intervals in which the state remains continuously unchanged are searched, and the start position, end position, and duration of each continuous interval are calculated.

[0071] In practical applications, this continuous interval search algorithm is used for: (1) In the longest sound duration, / s / , / z / and other continuous sound production tests, search for the longest continuous effective sound production interval and use its duration as the sound production ability index; (2) In the continuous counting task, the continuity of the effective vocal range is analyzed in combination with the allowed short pause conditions; (3) In the breathing-voice coordination analysis, search for continuous silent intervals and determine whether their duration exceeds the preset threshold in order to identify abnormal ventilation behavior.

[0072] The continuous interval search process does not rely on complex models, but is based on time series traversal and interval length calculation based on frame-level determination results. The duration of the continuous interval can be expressed as T=N×Δt; where N is the number of consecutive frames and Δt is the frame shift time.

[0073] In addition, the parameter configuration process can be achieved through the result display process of the software interface, and is not limited here.

[0074] Figure 2 A flowchart of another data processing method provided in the embodiments of this application is shown below. Figure 2 As shown, this step includes: S21: Voice data acquisition and preprocessing; S22: Voice data frame processing; S23: Extract acoustic features from the framed speech data; S24: Determine the vocal energy threshold; S25: Identification of the maximum effective vocal range within a continuous effective vocal range; S26: Determine the actual vocal performance indicators; S27: Determine the breath-voice coordination analysis; S28: Output of task-related metrics.

[0075] The foregoing has described in detail various embodiments corresponding to the data processing method. Based on this, this application also discloses a data processing apparatus corresponding to the above-described method. Figure 3 This is a structural diagram of a data processing apparatus provided in an embodiment of this application. Figure 3 As shown, the data processing device includes: The acquisition module 11 is used to acquire acoustic feature data obtained by feature extraction processing of the speech data to be tested and the corresponding current breathing and vocalization function task; wherein, the acoustic feature data includes at least energy feature data and fundamental frequency feature data; Processing module 12 is used to process acoustic feature data based on the task characteristics and vocal energy threshold of the current breathing and vocalization function task to obtain the corresponding effective vocalization range; wherein, the vocal energy threshold is pre-processed based on the current environmental noise energy feature data of the speech data. Module 13 is used to determine the vocal capacity index and the breathing-vocal coordination result based on the effective vocal range.

[0076] Since the embodiments of the device part correspond to the embodiments described above, please refer to the embodiments described in the method part for the embodiments of the device part, and will not be repeated here.

[0077] For a description of the data processing apparatus provided in this application, please refer to the above method embodiments. This application will not repeat the description here, as it has the same beneficial effects as the above data processing method.

[0078] Figure 4 A structural diagram of a breathing and vocalization ability determination system provided in an embodiment of this application is shown below. Figure 4 As shown, the system includes: Memory 21 is used to store computer programs; Processor 22 is used to implement the steps of a data processing method when executing a computer program.

[0079] The breathing and vocalization ability determination system provided in this embodiment can include, but is not limited to, smartphones, tablets, laptops, or desktop computers.

[0080] The processor 22 may include one or more processing cores, such as a quad-core processor or an octa-core processor. The processor 22 may be implemented using at least one of the following hardware forms: Digital Signal Processor (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 22 may also include a main processor and a coprocessor. The main processor, also known as the Central Processing Unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 22 may integrate a Graphics Processing Unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, the processor 22 may also include an Artificial Intelligence (AI) processor, which handles computational operations related to machine learning.

[0081] The memory 21 may include one or more computer-readable storage media, which may be non-transitory. The memory 21 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In this embodiment, the memory 21 is used to store at least the following computer program 211, which, after being loaded and executed by the processor 22, is capable of implementing the relevant steps of the data processing method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 21 may also include an operating system 212 and data 213, etc., and the storage method may be temporary storage or permanent storage. The operating system 212 may include Windows, Unix, Linux, etc. The data 213 may include, but is not limited to, the data involved in the data processing method, etc.

[0082] In some embodiments, the breathing and vocalization ability determination system may further include a display screen 23, an input / output interface 24, a communication interface 25, a power supply 26, and a communication bus 27.

[0083] Those skilled in the field can understand, Figure 3 The structure shown does not constitute a limitation on the system for determining breathing and vocalization capabilities and may include more or fewer components than illustrated.

[0084] The processor 22 implements the data processing method provided in any of the above embodiments by calling instructions stored in the memory 21.

[0085] For a description of the breathing and vocalization ability determination system provided in this application, please refer to the above method embodiments. This application will not repeat the description here, as it has the same beneficial effects as the above data processing method.

[0086] Furthermore, this application also provides a computer-readable storage medium storing a computer program, which, when executed by processor 22, implements the steps of the data processing method described above.

[0087] It is understood that if the methods in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0088] For a description of the computer-readable storage medium provided in this application, please refer to the above method embodiments. This application will not repeat the description here, as it has the same beneficial effects as the above data processing method.

[0089] The foregoing has provided a detailed description of a data processing method, a breathing and vocalization ability determination system, and a medium provided in this application. The various embodiments in the specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of this application.

[0090] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

Claims

1. A data processing method, characterized in that, include: Acquire acoustic feature data and the corresponding current breathing and vocalization function task obtained by feature extraction processing of the speech data to be tested; wherein, the acoustic feature data includes at least energy feature data and fundamental frequency feature data; Based on the task characteristics and vocal energy threshold of the current breathing and vocalization function task, the acoustic feature data is processed to obtain the corresponding effective vocalization range; wherein, the vocal energy threshold is pre-processed based on the current environmental noise energy feature data of the speech data. The vocal capacity index and the results of breath-voice coordination are determined based on the effective vocal range.

2. The data processing method according to claim 1, characterized in that, When the current breathing and vocalization task is a silent friction phonation task, the acoustic feature data is processed based on the task characteristics and vocal energy threshold of the current breathing and vocalization task to obtain the corresponding effective vocalization range, including: The data type of acoustic feature data is determined based on the task characteristics of the silent friction sound generation function task; wherein, the data type includes the data type of the energy feature data; Filter the energy feature data that exceeds the vocal energy threshold to select a first target energy feature data. If the time corresponding to each first target energy feature data is continuous and the continuous time is greater than the first preset continuous time, the sound emission interval corresponding to each first target energy feature data is taken as the effective sound emission interval.

3. The data processing method according to claim 1, characterized in that, When the current breathing and vocalization task is a continuous vocalization task, the acoustic feature data is processed based on the task characteristics and vocal energy threshold of the current breathing and vocalization task to obtain the corresponding effective vocalization range, including: The data type of acoustic feature data is determined based on the task characteristics of the continuous sound emission function task; wherein, the data type includes the data type of energy feature data and the data type of fundamental frequency feature data; The corresponding preset range of fundamental frequency is determined in advance based on the subject's age parameter corresponding to the voice data; Filter the energy feature data that exceeds the vocal energy threshold to select a first target energy feature data. The corresponding first target fundamental frequency feature data is determined based on the target speech data corresponding to the first target energy feature data; Select second target base frequency feature data that fall within the corresponding preset range of base frequency from each first target base frequency feature data; When the time corresponding to each second target fundamental frequency characteristic data is continuous and the continuous time is greater than the second preset continuous time, the sound range corresponding to each second target fundamental frequency characteristic data is taken as the effective sound range.

4. The data processing method according to claim 1, characterized in that, When the current breathing and vocalization function task is a continuous counting function task, the acoustic feature data is processed based on the task characteristics and vocal energy threshold of the current breathing and vocalization function task to obtain the corresponding effective vocalization range, including: The data type of acoustic feature data is determined based on the task characteristics of the continuous counting function; wherein, the data type includes the data type of energy feature data and the data type of fundamental frequency feature data; Filter the energy feature data that exceeds the vocal energy threshold to select a first target energy feature data. When the time corresponding to the first target energy feature data is continuous and the continuous time is greater than the third preset continuous time, the speech data corresponding to the first target energy feature data is aggregated to obtain a speech segment. Each sound segment is labeled as a discrete sound event, and the time position corresponding to each discrete sound event is marked to obtain the event label; The event markers are sorted to obtain a sequence of vocal events; If the time interval between adjacent sound events in the sound event sequence is less than a preset time interval, then the counting process corresponding to the sound event sequence is determined to be the initial counting process. The corresponding preset range of fundamental frequency is determined based on the subject's age parameter corresponding to the voice data of the initial counting process; Select third target base frequency feature data that falls within the corresponding preset range of base frequency from the base frequency feature data during the initial counting process; When the time corresponding to each third target fundamental frequency characteristic data is continuous and the continuous time is greater than the third preset continuous time, the sound interval corresponding to each third target fundamental frequency characteristic data is taken as the effective sound interval.

5. The data processing method according to any one of claims 1 to 4, characterized in that, Vocal performance indicators are determined based on the effective vocal range, including: Mark the time axis corresponding to the sound frames within the effective sound range to determine the continuous effective sound range; The maximum effective vocal range is obtained by extracting the continuous effective vocal range based on the start and end times and duration of the vocalization. Within the maximum effective vocal range, the actual vocal ability index is obtained by processing the speech data within the maximum effective vocal range according to the preset vocal ability indexes under the breathing vocal function task.

6. The data processing method according to claim 5, characterized in that, The maximum effective vocal range is obtained by extracting the continuous effective vocal range based on the start and end times and duration of the vocalization, including: Within the continuous effective sound emission interval, if the corresponding actual start and end time is greater than or equal to the fourth preset continuous time, and the number of events under the corresponding duration is greater than or equal to the preset number of events, then the continuous effective sound emission interval to which the corresponding actual start and end time is greater than or equal to the fourth preset continuous time, and the number of events under the corresponding duration is greater than or equal to the preset number of events, shall be taken as the maximum effective sound emission interval.

7. The data processing method according to claim 5, characterized in that, The process of determining the respiratory-vocal coordination result includes: Determine the target duration of the silent zone within the maximum effective vocal range; If the duration of the target exceeds the fifth preset continuous time, it is determined that there is a pause in speech or a gap in speech within the maximum effective vocal range, and the result of the breathing-voice coordination is insufficient. If no silent interval is detected, or if the target duration of the silent interval is less than or equal to the fifth preset continuous time, then the breathing and vocalization coordination result is determined to be good.

8. The data processing method according to claim 1, characterized in that, The process of determining the sound energy threshold includes: Obtain the current ambient noise data of the location where the voice data is located; Multiple current environmental noise energy feature data are obtained by feature extraction based on the current environmental noise data; The average environmental noise energy characteristic data is obtained by averaging multiple current environmental noise energy characteristic data. The sound energy threshold is determined based on the average environmental noise energy characteristic data. Correspondingly, determining the sound energy threshold based on the average environmental noise energy characteristic data includes: Obtain a first empirical coefficient and a second empirical coefficient; sum the average environmental noise energy characteristic data and the first empirical coefficient to obtain the sound energy threshold; Alternatively, obtain a second empirical coefficient; process multiple current environmental noise energy characteristic data by standard deviation to obtain standard environmental noise energy characteristic data; obtain first environmental noise energy characteristic data based on the second empirical coefficient and the standard environmental noise energy characteristic data; and sum the average environmental noise energy characteristic data and the first environmental noise energy characteristic data to obtain the sound energy threshold.

9. A system for determining respiratory vocalization ability, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the data processing method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the data processing method as described in any one of claims 1 to 8.