A method and apparatus for assessing speech expression ability in neuropsychiatric diseases
By combining large language models and natural language processing technology with speech recognition and multi-dimensional indicator evaluation, the problems of subjective bias and low efficiency in the assessment of speech expression ability in neuropsychiatric diseases have been solved, achieving efficient and accurate assessment results and providing a basis for disease diagnosis and rehabilitation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- THE SECOND XIANGYA HOSPITAL OF CENT SOUTH UNIV
- Filing Date
- 2026-01-05
- Publication Date
- 2026-05-12
AI Technical Summary
Existing methods for assessing the speech expression ability of neuropsychiatric patients suffer from subjective bias, inefficiency, high cost, and difficulty in comprehensively analyzing deep speech features. Existing automated solutions are not applicable to the assessment of patients with neuropsychiatric diseases.
By employing large language models and natural language processing technology, combined with speech recognition, semantic calibration and sentence segmentation, text error correction and multi-dimensional indicator evaluation, and guiding patients to tell stories through visual stimulation materials, and utilizing multi-model validation and human review, we can achieve fully automated and accurate assessment of speech expression ability.
It achieves objective, accurate, and comprehensive assessment of verbal expression ability, eliminates subjective bias, improves assessment efficiency, reduces costs, and enables in-depth analysis of various aspects of verbal expression ability, providing a basis for disease diagnosis and rehabilitation.
Abstract
Description
Technical Field
[0001] This application pertains to the medical field, specifically to a method and apparatus for assessing verbal expression ability in the field of neuropsychiatry. Background Technology
[0002] Schizophrenia, bipolar disorder, dementia, and aphasia are major neuropsychiatric disorders that seriously endanger human health. These diseases are often accompanied by significant impairment in verbal expression, and due to differences in neuropathological mechanisms, the specific patterns of impairment and their cognitive roots vary. Accurate assessment of verbal expression ability is of great value for analyzing the causes of verbal expression impairment, revealing the core neuropsychopathology of the disease, assisting in differential diagnosis, monitoring treatment efficacy, and guiding rehabilitation.
[0003] The current main assessment method is manual assessment, which has insurmountable bottlenecks: First, it is plagued by subjective bias, and it is difficult to unify the interpretation standards among different assessors or even the same assessor at different times, which seriously undermines the objectivity and repeatability of the results; second, the process is highly dependent on scarce clinical experts, resulting in low assessment efficiency and high costs, making it difficult to popularize in primary healthcare and large-scale screening; most importantly, human auditory perception is difficult to refine and deconstruct massive amounts of language data in a multidimensional way, causing assessments to mostly stay at the level of superficial acoustic features such as speech rate and pauses or simple vocabulary statistics, while the analysis of deep speech features such as semantic coherence, logical structure, information density and narrative efficiency, which are crucial to revealing the cognitive roots of speech ability impairment, is seriously insufficient or can only be described in a vague qualitative manner.
[0004] Several solutions exist for automated analysis of verbal expression ability, but most focus on assessing "eloquence" in general scenarios. The indicators (such as fluency, breadth and depth of thought, and persuasiveness) and models used in these solutions are designed to measure the upper limit of verbal expression ability in healthy individuals and to optimize verbal expression. These are fundamentally different from pathological speech disorders caused by diseases. Furthermore, they lack sufficient analysis of indicators (such as semantic coherence and circuitous expression) that reflect the cognitive roots of verbal expression impairment (such as flight of ideas, thought disorder, comprehension difficulties, and semantic retrieval difficulties), making them unsuitable for the clinical assessment of patients with the aforementioned neuropsychiatric disorders. In addition, traditional automated machine assessment methods can only perform mechanical mapping based on rules and cannot understand semantics, leading to potential scoring biases. Therefore, there is an urgent need in this field for a solution that can comprehensively overcome these shortcomings. Summary of the Invention
[0005] The purpose of this application is to provide a method and apparatus for assessing the verbal expression ability of neuropsychiatric disorders, which can achieve objective, accurate, efficient and comprehensive assessment.
[0006] The technical solution provided in this application is as follows:
[0007] In a first aspect, this application provides a method for assessing verbal expression ability in neuropsychiatric disorders, comprising:
[0008] Stimulus presentation and data acquisition: Visual stimuli and speech generation instructions are presented to the assessment subjects; speech data of the assessment subjects' narration based on the visual stimuli and speech generation instructions is acquired;
[0009] Speech data conversion and preliminary sentence segmentation: The speech data is processed, including: converting the speech data into text using automatic speech recognition technology, obtaining the start and end time information of each syllable, performing preliminary sentence segmentation on the text based on speech flow and semantics, and obtaining preliminary text sentence segmentation results;
[0010] Semantic calibration of sentence segmentation: The initial text sentence segmentation results are calibrated using a large language model to achieve accurate text sentence segmentation based entirely on semantics. The prompt words of the large language model include role setting, task background setting, task goal setting, execution step setting, constraint setting and output requirement setting.
[0011] Text correction: Natural language processing techniques are used to correct errors in the semantically calibrated sentence segmented text to obtain the final text;
[0012] Multi-dimensional indicator evaluation: Based on the data processed by speech data conversion and preliminary sentence segmentation, semantic calibration sentence segmentation and text error correction, natural language processing technology and large language model are used to evaluate one or more of the following indicators: fluency indicators, micro-linguistic indicators and macro-linguistic indicators.
[0013] Robustness verification: The robustness verification of the indicator results analyzed using large language models in the multi-dimensional indicator evaluation steps is carried out. The final results are output through independent evaluation by at least two independent large language models, iterative reflection, or manual review.
[0014] Output results: Outputs the final evaluation results and reliability information for multiple dimensions of indicators.
[0015] In one possible implementation, the fluency indicators in the multi-dimensional indicator evaluation step include speech rate, pronunciation speed, silent pause features, phrase flow and long speech flow features, and vocal pause features. Silent pauses are defined as pauses at non-semantic boundaries whose duration exceeds a first duration threshold or pauses at semantic boundaries whose duration exceeds a second duration threshold. Semantic boundaries are determined by the positions of punctuation marks in the text. Phrase flows are speech segments segmented by silent pauses whose duration exceeds the first duration threshold but is less than or equal to the second duration threshold. Long speech flows are speech segments segmented by silent pauses whose duration exceeds the second duration threshold. The average duration of phrase flows = total duration of phrase flow segments / number of phrase flow segments. The average duration of long speech flows = total duration of long speech flow segments / number of long speech flow segments. Vocal pauses include filler word pauses, utterance repetitions, utterance omissions, and utterance corrections, identified through a large language model combining text and speech data. Filler word pauses require that the syllable duration exceeds two standard deviations of the individual's average pronunciation speed or that there is silence greater than the first duration threshold after the filler word.
[0016] In one possible implementation, the micro-linguistic indicators mentioned in the multi-dimensional indicator evaluation step include total word count, total discourse count, word diversity, average discourse length, syntactic complexity, and grammatical accuracy. Specifically, when calculating the total word count, high-frequency words in the visual stimulus material are designated as segmentation hot words, and the effective vocabulary of non-punctuation and pause filler words is counted. Word diversity is calculated using both basic diversity and sliding window diversity. Syntactic complexity is determined by part-of-speech, syntactic tags, and semantic dependency tags to classify simple and complex sentences. Complex sentences include types such as "ba" sentences and "bei" sentences, and are measured by the proportion of complex sentences and average syntactic depth. Grammatical accuracy is achieved by identifying grammatical errors through a large language model, and the number of errors and the proportion of erroneous sentences are counted.
[0017] In one possible implementation, the macro-linguistic indicators in the multi-dimensional indicator evaluation step include topic accuracy and global coherence, event accuracy, story grammar and plot level, local coherence, topic deviation, and speech information rate. Specifically: topic accuracy and global coherence, event accuracy, and story grammar and plot level are evaluated using a large language model combined with corresponding standard information; local coherence is achieved by converting sentences into dense vectors using vector embedding technology, measured by the cosine similarity of adjacent sentence vectors. A cosine similarity greater than or equal to a first similarity threshold and less than a second similarity threshold is defined as semantic incoherence, and a cosine similarity less than the first similarity threshold is defined as semantic breakage; topic deviation is determined by the cosine similarity between sentence vectors and standard story topic vectors. A similarity greater than or equal to the first similarity threshold and less than the second similarity threshold indicates moderate deviation, and a similarity less than the first similarity threshold indicates severe deviation. The first similarity threshold is less than the second similarity threshold, and both are determined empirically.
[0018] In one possible implementation, the calculation of local coherence and topic deviation in the multi-dimensional index evaluation step is based on the final text after referential resolution by natural language processing technology, and the vector embedding technology uses Qwen or BERT models to implement semantic encoding.
[0019] In one possible implementation, the robustness verification step, in addition to topic accuracy, global coherence, and vocal pauses, involves the following verification method: A core large language model and an auxiliary verification large language model are used for independent dual-model evaluation, calculating the intra-group correlation coefficient or Kappa coefficient. When the coefficient is greater than or equal to a first correlation threshold, the core large language model result is output. When the coefficient is greater than a second correlation threshold but less than the first correlation threshold, up to N iterations of reflection are initiated, providing both models with critical reflection prompts containing the other's evaluation results and supporting evidence. If the model still fails to meet the standard after iteration and the highest coefficient during iteration is less than the second correlation threshold, manual review is initiated. The first correlation threshold is greater than the second correlation threshold, and both are determined empirically. The maximum number of iterations N is an empirical value. Topic accuracy and global coherence require complete consistency between the results of the two models to be considered robust and reliable. Vocal pauses require a difference of less than 4 in the total number of recognitions by the two models to be considered robust and reliable.
[0020] In one possible implementation, all evaluation tasks using a large language model in the multi-dimensional indicator evaluation step have prompts that include role setting, task goal setting, constraint setting, and output format setting, and supplement them with exclusive evaluation rules and standard information for different indicators.
[0021] Secondly, this application provides a device for assessing the verbal expression ability of neuropsychiatric disorders, comprising:
[0022] Stimulus presentation and data acquisition module: presents visual stimulus materials and speech generation instructions to the assessment subject; and acquires speech data of the assessment subject's narration based on the visual stimulus materials and speech generation instructions.
[0023] The data processing module is used to perform the following operations:
[0024] Speech data conversion and preliminary sentence segmentation: The speech data is processed, including: converting the speech data into text using automatic speech recognition technology and obtaining the start and end time information of each syllable; performing preliminary sentence segmentation on the text based on speech flow and semantics to obtain preliminary text sentence segmentation results;
[0025] Semantic calibration of sentence segmentation: The initial text sentence segmentation results are calibrated using a large language model to achieve accurate text sentence segmentation based entirely on semantics. The prompt words of the large language model include role setting, task background setting, task goal setting, execution step setting, constraint setting and output requirement setting.
[0026] Text correction: Natural language processing techniques are used to correct errors in the semantically calibrated and segmented text to obtain the final text;
[0027] Multi-dimensional indicator evaluation: Based on the data processed by speech data conversion and preliminary sentence segmentation, semantic calibration sentence segmentation and text error correction, natural language processing technology and large language model are used to evaluate one or more of the following indicators: fluency indicators, micro-linguistic indicators and macro-linguistic indicators.
[0028] Robustness verification: The robustness verification of the indicator results analyzed using large language models in the multi-dimensional indicator evaluation steps is carried out. The final results are output through independent evaluation by at least two independent large language models, iterative reflection, or manual review.
[0029] The results presentation module is used to output the final evaluation results and reliability information of indicators across multiple dimensions.
[0030] The device uses the method described in the first aspect above to assess the verbal expression ability of neuropsychiatric patients.
[0031] Thirdly, this application provides a computer-readable storage medium storing a computer program that, when run on an electronic device, causes the electronic device to perform the method described above.
[0032] Fourthly, this application provides a computer program product, including a computer program that, when run on an electronic device, causes the electronic device to perform the method described above.
[0033] The specific implementation methods of the second to fourth aspects of this application can refer to the implementation methods of the first aspect, and will not be elaborated here.
[0034] Beneficial effects:
[0035] 1. Objective and accurate quantification: Through fully automated processes and deep semantic analysis of large models, the subjective bias of manual evaluation is completely eliminated. It can accurately measure deep language features that were previously difficult to quantify (such as semantic coherence and narrative structure), making the evaluation results highly objective and repeatable.
[0036] 2. High efficiency and economy: The automation of the assessment process greatly improves assessment efficiency, significantly reduces reliance on senior experts and time and manpower costs, making large-scale and rapid screening possible, which is conducive to the application and promotion of the technology in primary healthcare institutions.
[0037] 3. Comprehensive and in-depth analysis: Based on the cognitive roots of speech expression impairment in neuropsychiatric diseases, a comprehensive and multi-level analytical framework is systematically constructed, ranging from surface fluency to deep semantic logic. This framework can more precisely characterize all aspects of speech expression ability, thereby providing a basis for analyzing the causes of speech expression impairment, revealing the core neuropsychopathology of the disease, assisting in differential diagnosis, monitoring treatment efficacy, and providing rehabilitation guidance.
[0038] 4. Reliable and trustworthy results: Exclusive prompt words were designed for the field of spontaneous language expression ability assessment in open scenarios, enabling large language models to output accurate assessment results. At the same time, a robust verification mechanism based on iterative reflection of two large language models effectively ensures the credibility of the AI model's assessment results on key indicators, solves the core pain point of AI application in serious medical scenarios, and provides a solid guarantee for the clinical reference value of the assessment results.
[0039] 5. Technological Integration and Innovation: By combining large language models, traditional natural language processing techniques, and vector embedding techniques, the strengths of each are combined to form a complementary technical solution, enabling accurate analysis of complex speech features. Detailed Implementation
[0040] To enable those skilled in the art to better understand the present application, the technical solution of the present application will be further described in detail below with reference to the embodiments of the present application.
[0041] It should be noted that the terms "first," "second," etc., used in this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be used interchangeably where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0042] The specific implementation of this application will now be described.
[0043] Example 1:
[0044] This application provides a method for assessing verbal expression ability in neuropsychiatric disorders, including:
[0045] S1. Stimulus Presentation and Data Acquisition: Present visual stimuli and speech generation instructions to the assessment subject; collect speech data of the assessment subject's narration based on the visual stimuli and speech generation instructions.
[0046] In some embodiments, the visual stimulus material is a set of story pictures with multiple events, multiple plots, and coherent content, used for picture description.
[0047] Picture description, as a verbal evoked task, can elicit spontaneous speech rich in information and form an objective frame of reference based on a unified stimulus, thus providing a comparable basis for evaluations by different individuals at different times. It is an efficient paradigm.
[0048] It should be understood that, in addition to story pictures, the visual stimuli can also employ different paradigms such as thematic narratives as verbal evoked tasks.
[0049] The visual stimulus material is standardized. The standard information derived from the visual stimulus material includes a standard story theme, a standard story event summary, a standard story plot summary, and a target vocabulary list, which covers relevant nouns, verbs, adjectives, and synonyms and near-synonyms.
[0050] The speech data (such as the story content being told) narrated by the assessment subject based on the visual stimulus material and speech generation instructions is collected (recorded) through a data acquisition module (such as a high-fidelity noise reduction recording device).
[0051] S2. Speech data conversion and preliminary sentence segmentation: The speech data is processed, including: converting the acquired speech data into text using automatic speech recognition technology, and obtaining the start and end time information of each syllable in the recording file; performing preliminary sentence segmentation on the text based on speech flow and semantics to obtain preliminary text sentence segmentation results.
[0052] In some embodiments, the acquisition of speech data is converted into text using automatic speech recognition technology, including: setting high-frequency words (such as the names of key characters, key action words, and key scene words in story pictures) corresponding to the visual stimulus material as speech recognition hot words and importing them into the hot word library of the automatic speech recognition system; and combining the hot word library with the acquisition of speech data to convert it into text using automatic speech recognition technology.
[0053] In this step, based on a hot word library, the acquired speech data is converted into text using automatic speech recognition technology. The hot word library will prioritize matching the core words involved in the evaluation object's description, reducing misidentification caused by accents, noise, or polysemy (such as avoiding misidentifying "nurse" as "neglect" or "injection" as "score"), ensuring the accuracy of core word recognition, reducing the redundant workload of subsequent error correction steps, and preventing the error correction algorithm from mistakenly modifying core words, further improving text quality and laying the foundation for subsequent indicator calculations.
[0054] Speech flow refers to the rhythm of speech flow during the evaluation of the subject's narration. Based on the "start time information of each syllable" obtained in step S2—such as natural pauses (short pauses without semantic boundaries) and changes in speech rate—it serves as the formal basis for sentence segmentation. For example, the logic for segmentation based on speech flow includes: identifying pauses (without semantic boundaries) with a duration ≥ a set time threshold (e.g., 2200ms) and nodes where the speech rate significantly slows down, as candidate points for sentence segmentation.
[0055] Semantics refers to the inherent logical meaning of the text itself, such as the collocation relationships between words, the completeness of phrases, and the coherence of the intended meaning, serving as the basis for sentence segmentation. For example, the logic of semantic-based segmentation includes: identifying word collocation relationships in the text (such as core verbs and noun collocations) and the boundaries of simple semantic units to assist in determining the position of clauses.
[0056] By combining speech flow and semantics, continuous text can be broken down into multiple preliminary sentence segments, avoiding excessively long segments that cross semantic boundaries or excessively short segments that are not semantically complete.
[0057] S3. Use a large language model to calibrate the initial text segmentation results to achieve accurate text segmentation based entirely on semantics.
[0058] The large language model prompts include character settings, task background settings, task goal settings, execution step settings, constraint settings, and output requirement settings.
[0059] In this application, when using a large language model to complete specific tasks (such as text segmentation calibration, spoken pause recognition, grammatical accuracy evaluation, etc.), the input to the model, in addition to the "original text to be analyzed" (such as the speech-to-text of the evaluation object), must also include the settings of 7 core modules to ensure that the model can accurately understand the task requirements and output results that meet expectations. These modules include:
[0060] Role setting: Clearly define the model's identity (e.g., "verbal assessment expert", "grammar correction specialist");
[0061] Task background setting: Describe the application scenario of the task (e.g., "clinical speech assessment of neuropsychiatric diseases", "open-ended spontaneous language analysis"): 3. Task objective setting: Clarify the specific tasks that the model should complete (e.g., "sentence segmentation based on semantic calibration", "identification of speech repetition", "assessment of grammatical error types");
[0062] Execution steps are set: inform the model of the specific process for completing the task (e.g., "first split the text, then check the semantic logic sentence by sentence", "first identify the filler words, then determine whether there is an audio pause").
[0063] Constraints are set: Define the output boundaries of the model (e.g., "segmentation is based solely on text semantics, without modifying the original expression" and "grammatical errors are strictly judged according to the definition, without omission or addition").
[0064] Output requirements: Define the output format of the model (e.g., "output the calibrated sentence text", "statistics on the number and percentage of each type of grammatical error");
[0065] By designing structured prompts, we ensure that the model accurately understands the task requirements, thus avoiding the situation where the output of a large language model fails to meet the accuracy and standardization requirements of clinical assessment due to "poor understanding." This is also one of the key technical details in this application that ensures the reliability of the assessment results.
[0066] S4. Use natural language processing technology to perform text error correction on the precise text sentences to obtain the final text.
[0067] This step uses natural language processing technology to perform text error correction on the precise text sentences, eliminating problems such as similar-looking characters and audio pauses, thereby further improving text quality.
[0068] S5. Based on the data processed in steps S2, S3 and S4, natural language processing techniques and large language models are used to evaluate multiple dimensions of speech expression indicators that are closely related to the causes of speech expression impairment in neuropsychiatric diseases. The dimensions include one or more of the following: fluency indicators, microlinguistic indicators and macrolinguistic indicators.
[0069] The fluency index focuses on the "fluency characteristics" of speech expression in patients with neuropsychiatric disorders, and can include five sub-indicators: speech rate, pronunciation speed, silent pause characteristics, characteristics of phrase flow and long flow, and vocal pause characteristics. The specific definitions of each sub-indicator are as follows:
[0070] Speech rate, which refers to the total number of syllables spoken by the subject per unit of time, reflects the overall pace of speech expression.
[0071] Speech rate indicates the efficiency of syllable output during the actual speech period (excluding silence and pauses), reflecting the effective rhythm of expression;
[0072] Silent pause characteristics include the number and frequency of silent pauses in speech, their cumulative duration and the proportion of the total duration, reflecting the interruption of expression;
[0073] The characteristics of phrase flow and long flow include the number and average duration of speech segments divided by silence pauses of different durations, reflecting the coherence of the speech flow;
[0074] The characteristics of vocal pauses include the number and frequency of pauses in speech that are accompanied by filler words, repetitions, or other vocal forms, reflecting the fluency of expression.
[0075] The calculation logic and core basis of each indicator:
[0076] The formula for calculating speech rate is: Speech rate = total number of syllables / total speech duration; the calculation is based on the "total number of syllables" obtained by speech recognition in step S2 + the "total speech duration" of the recording file.
[0077] The formula for calculating pronunciation speed is: Pronunciation speed = Total number of syllables / Pronunciation duration, Pronunciation duration = Total speech duration - Silence time. The calculation is based on: "Total number of syllables" and "Total speech duration" from step S2 + "Silence time" calculated based on the "syllable start and end time information" from S2.
[0078] Sub-indicators of the silence pause feature may include: number of silence pauses, total duration of silence pauses, and percentage of silence pauses (total duration / total speech duration). A silence pause is defined as a pause that meets any of the following conditions: ① The silence pause is defined as the start time of the following syllable minus the end time of the preceding syllable i > a first duration threshold (e.g., 200ms), and its position is located at a non-semantic boundary; ② The silence duration at a semantic boundary > a second duration threshold (e.g., 1000ms). The semantic boundary is determined by the positions of commas, periods, question marks, and exclamation marks in the text output in step S2.
[0079] Sub-indicators for phrase flow and long speech flow characteristics may include: number of phrase flow segments, number of long speech flow segments, average duration of phrase flow segments, and average duration of long speech flow segments. A phrase flow segment is defined as: all speech segments (expressive segments with medium-length pauses reflecting basic speech flow coherence) segmented by silent pauses with a duration greater than a first duration threshold (e.g., 200ms) and less than or equal to a second duration threshold (e.g., 1000ms). A long speech flow segment is defined as: all speech segments (expressive segments with long pauses reflecting sustained expressive ability) segmented by silent pauses with a duration greater than the second duration threshold. Silent pauses ≤ the first duration threshold are considered natural pauses within the speech flow, are not segmented, and are included in the current speech flow. The average duration of phrase flow segments is calculated as: Average duration of phrase flow segments = Total duration of phrase flow segments / Number of phrase flow segments; the average duration of long speech flow segments is calculated as: Average duration of long speech flow segments = Total duration of long speech flow segments / Number of long speech flow segments. The calculation is based on: the "syllable start and end time information" in step S2 (used to calculate the total duration of the segment and the number of segments).
[0080] The sub-indicators of the spoken pause feature can include: the number of spoken pauses and the frequency of spoken pauses (number of times / total speech duration). The types and definition criteria of spoken pauses include: ① Filler word pauses: The text contains words or phrases from a predefined list of filler words (such as "um," "ah," "that"), and satisfies the following conditions: the duration of any syllable of the word or phrase exceeds two standard deviations of the individual's average pronunciation speed, or there is a silence of >200ms after the filler word; ② Speech repetition: The phenomenon of continuous repetition of core words / phrases identified by the large language model; ③ Speech abandonment: The phenomenon of interruption before completion of expression identified by the large language model; ④ Speech correction: The phenomenon of supplementing or modifying expressed content identified by the large language model. The calculation logic for the number of spoken pauses and their frequency is as follows: Calculate the sum and frequency of the above four types of spoken pauses. The large language model prompts include role settings, task goal settings, evaluation indicator definitions, evaluation indicator examples and explanations, constraint settings, and output format settings. In other words, in addition to using the general framework described in step S3, the large language model prompt words also include additional evaluation index definitions, evaluation index cases and explanations based on the application scenarios (such as "word repetition in regular expressions should not be regarded as discourse repetition" and corresponding examples), to ensure that the recognition standards are consistent.
[0081] The core improvement in this step lies in addressing the challenge of identifying pauses in open, spontaneous language scenarios (such as "picture-based storytelling" by patients with neuropsychiatric disorders) that are unstructured and exhibit pathological expressive features. By leveraging the semantic understanding capabilities of a large language model and the acoustic features of the speech data, it identifies vocal pauses, including filler word pauses, repetitions, omissions, and corrections (based on the data obtained in step S3). This achieves a dual-dimensional "semantic + acoustic" recognition, covering common pathological vocal pauses in patients with neuropsychiatric disorders (such as omissions in schizophrenia patients and corrections in dementia patients), solving the problem of traditional methods only identifying filler words and missing core pathological features. It can integrate speech data (data obtained in step S2) and semantic information (based on the data obtained in step S3) to identify silent pauses, distinguishing between "normal pauses between sentences" (such as natural pauses at the end of sentences) and "pathological silent pauses" (such as long, non-semantic boundary pauses caused by interruptions in expression), avoiding misjudgments caused by traditional recognition methods based solely on speech duration, and improving the accuracy of silent pause indicators (frequency, duration percentage).
[0082] Among them, the micro-linguistic indicators focus on the "lexical and syntactic features" of speech expression in patients with neuropsychiatric disorders. They primarily reflect the ability to use vocabulary, construct sentences, and adhere to grammatical norms, and can include six sub-indicators: total word count, total number of discourses, word diversity, average discourse length, syntactic complexity, and grammatical accuracy assessed through natural language processing technology and / or large language models. The definitions of each sub-indicator are as follows:
[0083] Total word count represents the total number of effective words in the narrative text of the evaluated object (excluding punctuation and pause filler words), reflecting the productivity of verbal expression;
[0084] Total number of discourses represents the number of expressions in complete semantic sentences, reflecting the productivity of verbal expression;
[0085] Lexical diversity indicates the richness and repetition of vocabulary in a text, reflecting the flexibility of vocabulary reserves / use and the richness of linguistic output;
[0086] Average discourse length represents the average vocabulary size of a single sentence / discourse unit, reflecting the complexity of sentence construction;
[0087] Syntactic complexity represents the structural complexity (simple / complex sentence) and syntactic level depth of a sentence, reflecting the complexity of sentence construction and syntactic organization ability;
[0088] Grammatical accuracy refers to the degree to which grammatical rules are followed in verbal expression, reflecting the standardization of grammatical usage.
[0089] The calculation logic and core basis of each indicator are as follows:
[0090] The total word count is calculated as follows: High-frequency words from the visual stimulus material in S1 are set as hot words for word segmentation. After segmenting the text obtained in S4 using natural language processing (NLP) technology, the total word count is calculated, excluding punctuation and pause filler words. The calculation is based on: high-frequency words from S1 (hot words for word segmentation) + corrected text from S4 + NLP word segmentation technology.
[0091] The total number of discourses is calculated as follows: based on the text obtained in step S3, each period, question mark, or exclamation mark is treated as a discourse and the total number of discourses is calculated. The calculation is based on the text after semantic calibration in S3 plus the discourse boundaries defined by punctuation marks.
[0092] The word diversity calculation method includes two dimensions: ① Basic diversity: the number of different words within the total word count range; ② Sliding window diversity: based on the actual order of words in the text, with 50 words per analysis window and a moving window of 1 word length, the sum of the number of different words in all windows is calculated. This is then divided by (50 × number of windows). The calculation is based on the S4-corrected text (based on the total word count) + natural language processing sliding window analysis technology.
[0093] The method for calculating the average length of a discourse includes three dimensions: ① Dimension 1: Total word count / Total discourse count, reflecting the average vocabulary size of a single semantic sentence; ② Dimension 2: Total word count / Number of phrase stream segments, reflecting the average vocabulary size of phrase stream segments; ③ Dimension 3: Total word count / Number of long discourse stream segments, reflecting the average vocabulary size of long discourse stream segments. The calculation is based on: S4 corrected text (total word count) + S3 semantically calibrated text (total discourse count) + S2 syllable start and end time information (number of phrase streams / number of long discourse stream segments).
[0094] The syntactic complexity calculation method is as follows: Based on the S4-corrected data, natural language processing (NLP) techniques are used to obtain the part-of-speech (POS), syntactic tags, and semantic dependency tags for each word in each sentence. Based on the POS, syntactic tags, semantic dependency tags, and syntactic rules, a sentence is determined to be either a simple sentence or a complex sentence (complex sentences include: "ba" sentences, "bei" sentences, relative clauses, pivotal clauses, object complement clauses, serial verb constructions, etc.), and the syntactic depth (the average depth from the root node to each leaf node) is calculated. Syntactic complexity is measured by the number of complex sentences, the percentage of complex sentences in the total discourse, and the average syntactic depth. The calculation is based on: S4-corrected text + NLP POS / syntactic / semantic dependency annotation techniques.
[0095] The method for calculating grammatical accuracy is as follows: Based on the text after semantic calibration in step S3, natural language processing technology and a large language model are used to identify the grammatical errors to be evaluated. The large language model's prompts include character settings, task goal settings, evaluation content and definitions, scoring process settings, basic elements of the story content, evaluation indicator examples and explanations, constraint settings, and output format settings. In other words, in addition to using the general framework described in step S3, the large language model's prompts supplement the evaluation content and definitions (such as specific grammatical error types), scoring process settings, basic elements of the story content, and evaluation indicator examples and explanations—all tailored to the specific needs of grammatical evaluation. The grammatical errors to be evaluated can include one or more of the following 11 categories: misuse of word class, incorrect use of quantifiers, word order errors, incomplete components, redundant components, verb-object collocation errors, modifier-headword collocation errors, mixed sentence structures, unclear referents, missing grammatical function words, and roundabout expressions. The frequency of each type of grammatical error, the number of grammatically incorrect sentences, and the proportion of grammatically incorrect sentences in the total discourse are calculated as indicators of grammatical accuracy. The grammatical accuracy calculation is based on: S3 semantically calibrated text + natural language processing technology + a large language model with customized prompt words.
[0096] The core improvements in this step are: ① Setting high-frequency words corresponding to the visual stimulus material in S1 as hot words, and using natural language processing technology for word segmentation, part-of-speech tagging, syntactic tagging, and semantic role tagging, laying the foundation for accurate calculation of indicators such as total word count, total speech count, word diversity, average speech length, and syntactic complexity; ② Combining large language models and natural language processing technology, conducting specialized assessments of grammatical error types related to neuropsychiatric diseases, and using multi-technology fusion to assess grammatical accuracy, solving the pain point of traditional methods' difficulty in accurately identifying pathological grammatical problems, and achieving grammatical accuracy assessment; ③ The grammatical accuracy index is based on the semantic calibration text analysis in step S3 (avoiding the over-correction of errors masking real grammatical problems), while other indicators are based on the text analysis after error correction in step S4 (eliminating interference from speech recognition errors, speech disfluency, etc.), avoiding assessment bias caused by problems such as speech disfluency and inaccurate speech recognition, and hierarchical data adaptation, which can balance the authenticity and accuracy of the assessment and reduce bias.
[0097] Macro-linguistic indicators focus on the "overall semantics, narrative logic, and thematic relevance" of the speech expressions of patients with neuropsychiatric disorders, primarily reflecting narrative completeness, semantic coherence, and thematic fit. These can include six sub-indicators: thematic accuracy and global coherence, event accuracy, story grammar and plot level, as well as local coherence, thematic deviation, and speech information rate assessed through vector embedding techniques.
[0098] The specific definitions of each sub-indicator are as follows:
[0099] Thematic accuracy and overall coherence: Thematic accuracy refers to the degree to which the narrative content matches the standard story theme; overall coherence refers to the semantic and logical coherence of the entire narrative.
[0100] Event accuracy: This refers to the precision with which the core elements (subject, action, etc.) of each event in a standard story are described in the narrative;
[0101] Story grammar and plot level: Story grammar refers to the completeness of core grammatical components such as background, characters, and ending in a narrative; plot level refers to the degree of narrative structure of a single plot.
[0102] Local coherence: indicates the degree of semantic connection between adjacent sentences, reflecting the coherence of short-distance semantic logic;
[0103] Theme deviation: This indicates the degree to which the narrative deviates from the standard story theme, reflecting the narrative's ability to focus on the theme;
[0104] Verbal information rate: This refers to the output efficiency of target words related to the standard story in the narrative, reflecting the ability to effectively convey information.
[0105] The calculation logic and core basis of each sub-indicator:
[0106] All sub-indices of the macro-linguistic indicators are based on the data after text correction in step S4. Some require the use of natural language processing techniques to resolve pronouns. The specific calculation logic is as follows:
[0107] The calculation method for thematic accuracy and global coherence is as follows: The data and prompts after text correction in step S4 are input into the large language model, which then evaluates the indicators. The prompts in the large language model include character settings, task objective settings, standard story themes, evaluation indicator definitions and scoring rules, constraint settings, and output format settings. In other words, the prompts in the large language model follow the general framework described in step S3, but additionally include standard story themes, indicator definitions, and scoring rules. The scoring rules for thematic accuracy include: scoring based on the number of standard story themes, with 1 point awarded for each theme met, and the total score being the sum of all theme scores. The scoring rules for global coherence include: using a binary scoring method, 1 point is awarded for overall narrative logical coherence, and 0 points are awarded for incoherence.
[0108] The method for assessing event accuracy is as follows: Input the data and prompts after text correction in step S4 into the large language model to match each sentence with a corresponding standard story event. Evaluate the accuracy of the description from five dimensions: subject, action, scene, object, and internal reaction. Output the results, including: ① the number of accurate event descriptions (all five elements are correct), partially accurate (elements are incomplete but all existing elements are accurate), missing (no corresponding text for the event), and errors (any element is incorrect); ② the number of errors and missing elements in each of the five dimensions: subject, action, scene, object, and internal reaction. The prompts in the large language model include character settings, task objective settings, standard story event summary, evaluation process settings, scoring rule settings, constraint settings, and output format settings. In other words, the prompts in the large language model follow the general framework described in step S3, supplemented by a standard story event summary, evaluation process, and scoring rules. The accuracy assessment is based on the text after correction in S4 plus the standard story event summary.
[0109] The method for assessing story grammar and plot level is as follows: The data and prompts after text correction in step S4 are input into the large language model. For story grammar level, the large language model breaks down the original text into several sub-plots based on the standard story plot summary. It evaluates the completeness of various grammatical components (including background, characters, start event, internal reaction, plan, action, and ending—7 categories) in each sub-plot, outputting the sum of scores for each grammatical component across all sub-plots. For plot level, the large language model classifies and evaluates each plot. The classification results include complete plots, incomplete plots, reaction sequences, action sequences, or description sequences, outputting the quantity of each type. The prompts in the large language model include character settings, task goal settings, standard story plot summary, evaluation indicators and definitions, scoring rules, evaluation indicator examples and explanations, constraint settings, and output format settings. In other words, the prompts in the large language model, in addition to using the general framework described in step S3, also supplement the standard story plot summary, evaluation indicators and definitions, scoring rules, and evaluation indicator examples and explanations. The assessment of story grammar and plot level is based on the text after correction in S4 plus the standard story plot summary.
[0110] The method for calculating local coherence is as follows: Based on the data obtained in step S4, and after dereference resolution using natural language processing technology, each sentence is taken as a basic analysis unit. Vector embedding technology is used to transform each sentence into a dense vector (e.g., a 1024-dimensional dense vector). The cosine similarity between the current sentence and the previous sentence vector is calculated as a measure of local coherence. The average of all cosine similarities is calculated as the overall local coherence, and the standard deviation of all cosine similarities is calculated as the stability of local coherence. Semantic incoherence is defined as a first similarity threshold (e.g., 0.3) ≤ cosine similarity < second similarity threshold (e.g., 0.5), and semantic discontinuity is defined as a cosine similarity < first similarity threshold (e.g., 0.3). The number of semantic incoherences and semantic discontinuities and their proportions relative to the total number of utterances are calculated. Finally, the overall local coherence, local coherence stability, and the number and proportion of semantic incoherences and semantic discontinuities are output.
[0111] The method for calculating topic deviation is as follows: Based on the data obtained in step S4, and after dereference is resolved using natural language processing technology, each sentence in the standard story theme and narrative text is encoded into a vector. The cosine similarity between each sentence vector and all theme vectors is calculated (i.e., each sentence is used as the basic unit of analysis, converted into a vector, and then its cosine similarity with the preset standard story theme is calculated). If the similarity between a sentence and all themes is greater than or equal to the first similarity threshold (e.g., 0.3) and less than the second similarity threshold (e.g., 0.5), it is judged as moderate topic deviation; if the similarity is less than the first similarity threshold (e.g., 0.3), it is judged as severe topic deviation. The number and proportion of deviation utterances are output. This step, based on the data obtained in step S4 and the standard story theme, after dereference is resolved using natural language processing technology, uses vector embedding technology to evaluate and calculate the number and proportion of deviation utterances.
[0112] The method for calculating the speech information rate is as follows: ① Based on the content of the story pictures and the standard story theme, determine the target vocabulary. Target vocabulary includes nouns, verbs, adjectives, adverbs, conjunctions, numerals, measure words, and onomatopoeia determined according to the content of the story pictures and the standard story theme. The above vocabulary should, as far as possible, include synonyms and near-synonyms that conform to the vocabulary of the story pictures and the standard story theme; ② Calculate two dimensions: the percentage of target vocabulary in the total number of words, and the number of target words spoken per minute. The calculation basis for this step includes: the S4 corrected text (based on the total word count) + the target vocabulary list related to the standard story.
[0113] The core improvements in this step include: ① Utilizing the language comprehension capabilities of a large language model to assess topic accuracy, global coherence, event accuracy, and story grammar and plot level, addressing the scoring bias caused by the inability of traditional rule-based mechanical mapping to understand semantics; ② Implementing sentence semantic encoding through vector embedding technology, transforming semantics into computable vectors, calculating cosine similarity between sentences and between sentences and topics, and using cosine similarity to accurately quantify local coherence and topic deviation, replacing traditional fuzzy qualitative descriptions; ③ Analyzing all indicators using the data obtained in step S4 to avoid assessment bias caused by issues such as poor fluency in spoken language and inaccurate speech recognition, while reducing the additional token consumption of the large language model, balancing assessment accuracy and efficiency; ④ Indicator design aligns with the narrative characteristics of patients with neuropsychiatric disorders (e.g., topic deviation in schizophrenia patients, incomplete plot in dementia patients), providing targeted data support for revealing the pathological mechanisms of these diseases.
[0114] All tasks using the large language model in this application employ a cue word design of 'general framework + scenario-based supplement'. The general framework includes core modules such as role setting and task background setting (see step S3 for details). The scenario-based supplementary content is determined according to the specific evaluation indicator requirements, ensuring the adaptability of the cue words and the accuracy of the evaluation results. The cue words include various aspects such as role setting, task goal setting, standard content setting related to visual stimuli in S1, execution step setting, evaluation content definition, specific examples and explanations of evaluation indicators, constraint setting, and standard output format setting. Regarding the cue word content explaining the constraints and evaluation indicators, in addition to the benchmark rules and conditions (such as the principle of objectivity and rigidity), the specific large language model style and common errors in model responses should also be integrated according to the specific evaluation items. The standard output format setting should include the output of the evaluation score and the output of the scoring basis. This cue word design structure has undergone iterative optimization, and the evaluation results are highly consistent with the human evaluation results, exhibiting high consistency in the output results.
[0115] It should be understood that the calculation method of specific indicators can be adjusted according to the actual application scenario, such as the window length and step size for word diversity calculation, the basic analysis unit and transformation dimension size of vector embedding, and the threshold standard for silence pause duration.
[0116] S6. Perform robustness (reliability) verification on the evaluation results of the indicators analyzed using the large language model in step S5.
[0117] The verification includes: independently evaluating the same metric using at least two independent large language models based on the same prompt words, calculating the consistency of the evaluation results; determining whether to initiate an iterative reflection process or enter a manual review process based on the degree of consistency, and using the verified result as the final output. The robustness verification specifically includes:
[0118] Calculate the within-group correlation coefficient or Kappa coefficient of the assessment results of the two large language models;
[0119] If the coefficient is greater than or equal to the first relevant threshold, the evaluation result is considered robust and reliable, and the result of the core large language model is used as the final output.
[0120] If the coefficient is between the second correlation threshold and the first correlation threshold, a new prompt word is provided to each model, requiring it to combine the evaluation results and scoring criteria of the other model, critically reflect on the initial output result according to the scoring rules, re-output the result, and recalculate the consistency coefficient. This iterative process is repeated a preset number of times.
[0121] If the standard is still not met after iteration and the highest coefficient during the iteration is lower than the second correlation threshold, the evaluation of the corresponding indicator will enter the manual annotation queue and be output by professionals after calibration.
[0122] The first correlation threshold is greater than the second correlation threshold, and can be set based on experience.
[0123] In some embodiments, in addition to topic accuracy, global coherence, and vocal pauses, the verification method is as follows: Besides the core large language model, an independent large language model is used as an auxiliary verification model. Each indicator is independently evaluated based on the same prompts. Then, the consistency of the evaluation results of the two large language models is evaluated using the intragroup correlation coefficient (for continuous variables) or the Kappa coefficient (for non-continuous variables). A coefficient ≥ 0.8 indicates robust and reliable evaluation results, and the result of the core large language model is used as the final output. If the coefficient is 0.7 < coefficient < 0.8, new critical reflection prompts are provided to each large language model based on the first round of dialogue, informing it of the inconsistency in its evaluation results. It is required to compare the evaluation results and basis of the other model, re-critically evaluate its own evaluation process / results according to the evaluation rules, re-output the results, and recalculate the intragroup correlation coefficient or Kappa coefficient. The above iterative process is repeated three times. If any iteration meets the standard, the iteration terminates. If all three iterations fail to meet the standard, the output of the core large language model with the highest consistency is used as the result, and a special annotation indicating poor robustness and reliability is added. If the coefficient is <0.7, the evaluation of the corresponding indicator enters the manual annotation queue. The content provided to the manual reviewer includes the analyzed narrative text, the original scoring results, and the scoring basis. The results are calibrated by professional psychological assessors and output after calibration. For topic accuracy and global coherence, the robustness and reliability standard is that the results of the two large language models are completely consistent. If they are inconsistent in all three iterations, they also enter the manual review queue. The content provided to the manual reviewer includes the narrative text, the scoring results, and the scoring basis. For spoken pauses, the robustness and reliability standard is that the difference in the total number of spoken pauses calculated by the two large language models is <4. If the robustness and reliability standard cannot be met in all three iterations, they also enter the manual review queue. The content provided to the manual reviewer includes the narrative text, the scoring results, and the scoring basis. The scoring basis and scoring results will be selected based on the highest frequency of each spoken pause condition across all models during the iteration process.
[0124] S7. Output the final evaluation results and reliability information of multiple dimensions of indicators.
[0125] It should be understood that the above-mentioned consistency coefficient calculation method, consistency coefficient threshold (0.8, 0.7), maximum number of iterations and reflections (3 times) and other parameters can be adjusted according to the actual application scenario and model performance.
[0126] It should be understood that this invention does not rely on a specific large language model (such as DeepSeek, Qwen series, etc.), natural language processing tools (such as HanLP, LTP), speech recognition tools (such as Paraformer, Doubao, etc.) or vector embedding technology (such as Qwen, BERT, etc.), as long as it has sufficient data parsing capabilities.
[0127] It should be understood that the numbers S1 to S6 above are only used to distinguish and facilitate the expression of different steps, and do not necessarily constitute a restriction on the execution order between the steps.
[0128] Example 2:
[0129] This application provides an apparatus for assessing the verbal expression ability of neuropsychiatric disorders, comprising:
[0130] Stimulus presentation and data acquisition module: presents visual stimulus materials and speech generation instructions to the assessment subject; and acquires speech data of the assessment subject's narration based on the visual stimulus materials and speech generation instructions.
[0131] The data processing module is used to perform the following operations:
[0132] Speech data conversion and preliminary sentence segmentation: The speech data is processed, including: converting the speech data into text using automatic speech recognition technology and obtaining the start and end time information of each syllable; performing preliminary sentence segmentation on the text based on speech flow and semantics to obtain preliminary text sentence segmentation results;
[0133] Semantic calibration of sentence segmentation: The initial text sentence segmentation results are calibrated using a large language model to achieve accurate text sentence segmentation based entirely on semantics. The prompt words of the large language model include role setting, task background setting, task goal setting, execution step setting, constraint setting and output requirement setting.
[0134] Text correction: Natural language processing techniques are used to correct errors in the semantically calibrated sentence segmented text to obtain the final text;
[0135] Multi-dimensional indicator evaluation: Based on the data processed by speech data conversion and preliminary sentence segmentation, semantic calibration sentence segmentation and text error correction, natural language processing technology and large language model are used to evaluate one or more of the following indicators: fluency indicators, micro-linguistic indicators and macro-linguistic indicators.
[0136] Robustness verification: The robustness verification of the indicator results analyzed using large language models in the multi-dimensional indicator evaluation steps is carried out. The final results are output through independent evaluation by at least two independent large language models, iterative reflection, or manual review.
[0137] The results presentation module is used to output the final evaluation results and reliability information of indicators across multiple dimensions.
[0138] The device uses the method described in Embodiment 1 above to assess the verbal expression ability of neuropsychiatric disorders.
[0139] Example 3:
[0140] This embodiment provides a computer-readable storage medium storing a computer program. When the computer program is run on an electronic device, it causes the electronic device to perform the method described in Embodiment 1.
[0141] Example 4:
[0142] This embodiment provides a computer program product, including a computer program that, when run on an electronic device, causes the electronic device to perform the method described in Embodiment 1.
[0143] The specific implementation of the system, electronic device, computer-readable storage medium, and computer program product provided in this application can be referred to the specific embodiments of the above methods, and will not be repeated here.
[0144] Obviously, those skilled in the art should understand that the various units or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device, or fabricating them separately as individual integrated circuit modules, or fabricating multiple modules or steps into a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0145] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for assessing verbal expression ability in neuropsychiatric disorders, characterized in that, include: Stimulus presentation and data collection: Presenting visual stimuli and verbal generation instructions to the assessment subjects; Collect speech data of the assessment subject based on the visual stimulus material and speech generation instructions; Speech data conversion and preliminary sentence segmentation: The speech data is processed, including: converting the speech data into text using automatic speech recognition technology, obtaining the start and end time information of each syllable, performing preliminary sentence segmentation on the text based on speech flow and semantics, and obtaining preliminary text sentence segmentation results; Semantic calibration of sentence segmentation: The initial text sentence segmentation results are calibrated using a large language model to achieve accurate text sentence segmentation based entirely on semantics. The prompt words of the large language model include role setting, task background setting, task goal setting, execution step setting, constraint setting and output requirement setting. Text correction: Natural language processing techniques are used to correct errors in the semantically calibrated sentence segmented text to obtain the final text; Multi-dimensional evaluation indicators: Based on the data processed through speech data conversion and preliminary sentence segmentation, semantic calibration sentence segmentation, and text correction steps, natural language processing techniques and large language models are used to evaluate one or more of the following: fluency indicators, micro-linguistic indicators, and macro-linguistic indicators. The fluency indicators include speech rate, pronunciation speed, silent pause features, phrase flow and long flow features, and vocal pause features. The micro-linguistic indicators include total word count, total discourse count, word diversity, average discourse length, syntactic complexity, and grammatical accuracy. The macro-linguistic indicators include topic accuracy and global coherence, event accuracy, story grammar and plot level, local coherence, topic deviation, and speech information rate. Robustness verification: The robustness verification of the indicator results analyzed using large language models in the multi-dimensional indicator evaluation steps is carried out. The final results are output through independent evaluation by at least two independent large language models, iterative reflection, or manual review. Output results: Outputs the final evaluation results and reliability information for multiple dimensions of indicators.
2. The method according to claim 1, characterized in that, Silent pauses are defined as pauses at non-semantic boundaries whose duration exceeds a first duration threshold or pauses at semantic boundaries whose duration exceeds a second duration threshold. Semantic boundaries are determined by the positions of punctuation marks in the text. Phrase streams are speech segments segmented by silent pauses whose duration exceeds the first duration threshold but is less than or equal to the second duration threshold. Long speech streams are speech segments segmented by silent pauses whose duration exceeds the second duration threshold. The average duration of a phrase stream is equal to the total duration of phrase stream segments divided by the number of phrase stream segments. The average duration of a long speech stream is equal to the total duration of long speech stream segments divided by the number of long speech stream segments. Vocal pauses include filler word pauses, speech repetitions, speech omissions, and speech corrections. These are identified using a large language model that combines text and speech data. Filler word pauses require that the duration of the syllable exceeds two standard deviations or that there is silence greater than the first duration threshold after the filler word.
3. The method according to claim 1, characterized in that, When calculating the total word count, high-frequency words in the visual stimulus material are set as hot words for word segmentation, and the effective vocabulary of non-punctuation and pause filler words is counted. Word diversity is calculated using two dimensions: basic diversity and sliding window diversity. Syntactic complexity is determined by part of speech, syntactic tags and semantic dependency tags to distinguish between simple and complex sentences. Complex sentences include types such as "ba" sentences and "bei" sentences, and are measured by the proportion of complex sentences and average syntactic depth. Grammatical accuracy is assessed by identifying grammatical errors using a large language model and statistically analyzing the number of errors and the percentage of sentences with errors.
4. The method according to claim 1, characterized in that, Thematic accuracy and global coherence, event accuracy, and story grammar and plot level are assessed using a large language model combined with corresponding standard information. Local coherence is measured by converting sentences into dense vectors using vector embedding technology, with the cosine similarity between adjacent sentence vectors as the metric. A cosine similarity greater than or equal to the first similarity threshold and less than the second similarity threshold is defined as semantic incoherence, and a cosine similarity less than the first similarity threshold is defined as semantic discontinuity. Thematic deviation is determined by the cosine similarity between sentence vectors and standard story thematic vectors. A similarity greater than or equal to the first similarity threshold and less than the second similarity threshold is considered moderate deviation, and a similarity less than the first similarity threshold is considered severe deviation. The first similarity threshold is less than the second similarity threshold, and both are determined empirically.
5. The method according to claim 4, characterized in that, The calculation of local coherence and topic deviation in the multi-dimensional indicator evaluation steps are based on the final text after referential resolution is completed using natural language processing technology.
6. The method according to claim 1, characterized in that, In the robustness verification step, in addition to topic accuracy, global coherence, and vocal pauses, the verification method is as follows: a core large language model and an auxiliary verification large language model are used for independent evaluation of the two models, and the intra-group correlation coefficient or Kappa coefficient is calculated; when the coefficient is greater than or equal to the first correlation threshold, the core large language model result is output; when the coefficient is greater than the second correlation threshold but less than the first correlation threshold, up to N iterations of reflection are initiated, and critical reflection prompts containing the other's evaluation results and basis are provided to the two models; if the standard is still not met after iteration and the highest coefficient in the iteration is less than the second correlation threshold, manual review is initiated. The first relevance threshold is greater than the second relevance threshold, and both are determined based on experience; the maximum number of iterations N is an empirical value; the topic accuracy and global coherence are considered robust and reliable only if the results of the two models are completely consistent; the sound pauses are considered robust and reliable only if the difference in the total number of recognitions of the two models is less than 4.
7. The method according to claim 1, characterized in that, In the multi-dimensional indicator evaluation process, all evaluation tasks using the large language model have prompts that include role setting, task goal setting, constraint setting, and output format setting, and supplement them with exclusive evaluation rules and standard information for different indicators.
8. A device for assessing verbal expression ability in neuropsychiatric disorders, characterized in that, include: Stimulus presentation and data acquisition module: presents visual stimulus materials and speech generation instructions to the assessment subject; and acquires speech data of the assessment subject's narration based on the visual stimulus materials and speech generation instructions. The data processing module is used to perform the following operations: Speech data conversion and preliminary sentence segmentation: The speech data is processed, including: converting the speech data into text using automatic speech recognition technology and obtaining the start and end time information of each syllable; performing preliminary sentence segmentation on the text based on speech flow and semantics to obtain preliminary text sentence segmentation results; Semantic calibration of sentence segmentation: The initial text sentence segmentation results are calibrated using a large language model to achieve accurate text sentence segmentation based entirely on semantics. The prompt words of the large language model include role setting, task background setting, task goal setting, execution step setting, constraint setting and output requirement setting. Text correction: Natural language processing techniques are used to correct errors in the semantically calibrated sentence segmented text to obtain the final text; Multi-dimensional evaluation indicators: Based on the data processed through speech data conversion and preliminary sentence segmentation, semantic calibration sentence segmentation, and text correction steps, natural language processing techniques and large language models are used to evaluate one or more of the following: fluency indicators, micro-linguistic indicators, and macro-linguistic indicators. The fluency indicators include speech rate, pronunciation speed, silent pause features, phrase flow and long flow features, and vocal pause features. The micro-linguistic indicators include total word count, total discourse count, word diversity, average discourse length, syntactic complexity, and grammatical accuracy. The macro-linguistic indicators include topic accuracy and global coherence, event accuracy, story grammar and plot level, local coherence, topic deviation, and speech information rate. Robustness verification: The robustness verification of the indicator results analyzed using large language models in the multi-dimensional indicator evaluation steps is carried out. The final results are output through independent evaluation by at least two independent large language models, iterative reflection, or manual review. The results presentation module is used to output the final evaluation results and reliability information of indicators across multiple dimensions.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed on an electronic device, causes the electronic device to perform the method as described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is run on an electronic device, it causes the electronic device to perform the method as described in any one of claims 1 to 7.