Speech disorder assessment method and system based on multi-modal large model
By combining multimodal large models with cross-modal feature extraction and semantic alignment of audio, lip video and tongue ultrasound images, the problem of insufficient multimodal fusion and professionalism in the assessment of speech disorders in existing technologies is solved, and accurate assessment of speech disorders and personalized rehabilitation suggestions are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
- Filing Date
- 2026-04-13
- Publication Date
- 2026-07-14
AI Technical Summary
Existing methods for assessing speech disorders rely on single-modal analysis, lack multimodal data fusion, make it difficult to achieve accurate and personalized assessments, and have insufficient device portability and clinical usability, and the assessment results lack professional knowledge support.
Employing a multimodal large model, this study extracts and semantically aligns cross-modal features from audio signals, lip videos, and tongue ultrasound images. Combining this with the Frenchay scale criteria, it simulates clinical reasoning logic and introduces retrieval enhancement generation technology to generate a structured assessment report.
It enables precise extraction and comprehensive evaluation of multimodal data, improving the accuracy and professionalism of the assessment, meeting clinical needs, and providing authoritative rehabilitation advice.
Smart Images

Figure CN122392879A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, and more specifically, to a method and system for assessing speech disorders based on a multimodal large model. Background Technology
[0002] Articulation disorders are common speech impairments caused by neurological damage or diseases such as stroke and deafness, severely impacting patients' communication abilities and quality of life. Clinical studies show that without timely intervention, these disorders are prone to developing into chronic conditions, significantly reducing treatment response rates and increasing the difficulty of subsequent rehabilitation. Currently, the field of speech therapy faces two core challenges: firstly, the long training period and high professional development costs for qualified speech therapists have led to a significant global shortage of professionals, making it difficult for many patients to access timely and accessible diagnostic and treatment services; secondly, traditional treatment models are limited by time and space, making it difficult to meet patients' long-term, high-frequency rehabilitation needs, and manual assessment relies on subjective experience, resulting in insufficient consistency and accuracy. These problems highlight the urgent need to develop computer-assisted speech rehabilitation systems. Such systems need to achieve objective assessment and personalized intervention, becoming an important direction for clinical research. Early speech rehabilitation techniques mostly focused on single-modal analysis, such as extracting acoustic features like pitch, energy, and speech rate solely from audio signals, or capturing vocabulary usage preferences through text analysis to assist in diagnosing speech disorders. However, single-modal analysis cannot comprehensively reflect the motor state and pathological mechanisms of the speech organs, limiting its clinical value. With the development of medical imaging technology, Ultrasound Tongue Imaging (UTI), as a non-invasive and portable technology, can visualize the sagittal plane dynamics of the tongue in real time during continuous articulation, providing a key internal motion perspective for the assessment of articulation disorders, and has gradually become an important tool for the diagnosis of speech disorders.
[0003] In recent years, large multimodal models (MLLMs) have demonstrated enormous application potential in the medical field due to their powerful cross-modal information fusion and reasoning capabilities, particularly in medical image analysis, clinical decision support, and surgical planning. Through instruction fine-tuning techniques, MLLMs have been extended to vision-language tasks, exhibiting superior performance in image reasoning and video understanding. However, in the field of speech rehabilitation, existing computer-assisted speech therapy (CASLT) systems mostly employ retrieval-based dialogue models, which struggle to achieve natural and fluent human-computer interaction. Furthermore, limited by unimodal analysis, they neglect the complex relationships between audio, internal and external articulatory organs, and cannot provide accurate personalized feedback. Although MLLMs possess cross-modal processing capabilities, most are trained on general-domain data, lacking sufficient understanding of specialized medical modalities such as UTI (Understanding of Language-Based Inquiry) and struggling to interpret crucial information such as tongue movement trajectories. Moreover, the lack of cross-modal temporal information fusion mechanisms and the scarcity of large-scale labeled medical multimodal question-answering datasets further limit the generalization ability and clinical adaptability of these models in speech rehabilitation scenarios.
[0004] In existing technologies, there are several speech disorder assessment schemes based on speech, text, or images. For example, patent application CN2023107185050 discloses an artificial intelligence-based speech and language disorder assessment and rehabilitation training system, including an airflow detection device (containing a computer, whistle, balloon, dual cameras, and a housing) and a multi-module processing system. Parents or therapists generate a 3D virtual teaching assistant and assessment story script using keywords. Children interact with the virtual teaching assistant through speech and movement to complete the assessment. The virtual teaching assistant provides long-term companionship to collect data, record developmental "highlights," establish personalized norms to monitor progress, analyze interaction through a multimodal data processing model, automatically generate rehabilitation training courses, and provide targeted reinforcement training after assessment, achieving home-based, fun, and precise rehabilitation. Patent application CN201510050789.6 discloses a multi-parameter diagnostic and rehabilitation device and cloud rehabilitation system for speech, language, and hearing disorders based on acoustic, electrophysiological, and three-dimensional imaging modeling technologies. The device includes real-time measurement, main control, and real-time audiovisual feedback units. The real-time measurement unit collects multi-parameter data such as speech, voice, and articulation; the main control unit identifies the type of disorder and sets the treatment mode; and the real-time audiovisual feedback unit provides rehabilitation guidance and dynamic monitoring through various modern technologies. The cloud-based rehabilitation system connects to the diagnostic and treatment equipment, providing services such as dynamic diagnostic assessment, remote rehabilitation guidance, online training, and record management, forming a comprehensive rehabilitation system with multi-parameter modeling, precise diagnosis and treatment, and cloud support, thus promoting the intelligent development of rehabilitation medicine.
[0005] Analysis reveals the following main shortcomings in existing technologies: (1) Strong dependence on internal organ data collection: Existing multimodal speech rehabilitation / assessment methods often require direct collection of internal imaging data such as UTI, which has high requirements for equipment portability and operation, making it difficult to implement in scenarios such as home rehabilitation and grassroots screening.
[0006] (2) The assessment system is inconsistent with the Frenchay scale process: Many methods focus on a single modality or single organ (such as tongue movement) and lack the structured output of "organ sub-item - comprehensive judgment" consistent with the Frenchay scale, resulting in insufficient clinical usability.
[0007] (3) Insufficient multimodal synchronization and alignment mechanisms: There are significant differences in frame rate, latency and task boundaries between audio and video, and there is a lack of alignment for pronunciation events (such as phoneme boundaries / pronunciation task segments), which leads to cross-modal evidence mismatch and affects the credibility of reasoning.
[0008] (4) Insufficient professional knowledge support and lack of traceability: The conclusions and recommendations of the existing system rely heavily on the implicit knowledge of model parameters, lack real-time access and citation of guidelines / scale clauses / literature evidence, and are not clinically rigorous and interpretable. Summary of the Invention
[0009] The purpose of this invention is to overcome the shortcomings of the prior art and provide a speech disorder assessment method and system based on a multimodal large model.
[0010] According to a first aspect of the present invention, a method for assessing speech disorders based on a multimodal large model is provided. The method includes the following steps: For the target, acquire the corresponding multimodal data, which includes audio signals, lip video, and tongue ultrasound images; Cross-modal feature extraction and semantic alignment are performed on the multimodal data to map the extracted modal features to a unified semantic space aligned with the large model text embedding space, thereby obtaining a semantic vector sequence. Using the semantic vector sequence as input, the large model is used to simulate multi-level clinical reasoning logic to obtain preliminary assessment results of articulation disorders, including severity level and disorder type. Using the key information from the preliminary assessment conclusions, a retrieval strategy is set up to guide the large model in generating assessment reports and rehabilitation recommendations.
[0011] According to a second aspect of the present invention, a speech disorder assessment system based on a multimodal large model is provided. The system includes: Multimodal data acquisition module: used to acquire corresponding multimodal data for a target, including audio signals, lip video and tongue ultrasound images; Feature extraction and alignment module: used to perform cross-modal feature extraction and semantic alignment on the multimodal data, so as to map the extracted modal features to a unified semantic space aligned with the large model text embedding space, and obtain a semantic vector sequence; Preliminary assessment module: This module uses the semantic vector sequence as input and the large model to simulate multi-level clinical reasoning logic to obtain preliminary assessment results of articulation disorders. The preliminary assessment results include severity level and disorder type. Assessment report generation module: This module is used to set up a retrieval and query strategy using key information from the preliminary assessment conclusions, and to guide the large model to generate an assessment report and rehabilitation recommendations.
[0012] Compared with existing technologies, the advantages of this invention are that it constructs an intelligent assessment system for articulation disorders based on a multimodal large model and strictly follows the Frenchay clinical scale standards. This system can simulate and enhance the assessment process of professional speech therapists. The technical effects are mainly reflected in the following aspects: (1) Accurate extraction of cross-modal articulation features: Key features that can characterize articulation characteristics were effectively extracted from patient audio (acoustic signal), tongue ultrasound (internal articulation organ) and lip video (external organ movement), providing high-quality input for subsequent fine-tuning and alignment with the large model; (2) Alignment of non-textual modalities with the semantic space of large models: The extracted audio features, lip video features and tongue ultrasound image features generated from audio are transformed and aligned with the semantic understanding capabilities of large language models, so that the model can “read” these heterogeneous temporal data and understand multimodal pronunciation information. (3) Simulate the multi-level reasoning of clinical pathways: Set up a large model to imitate the structured diagnostic logic of clinicians. First, conduct independent sub-assessments and problem descriptions of the functions of core vocal organs such as lips, jaw, palate, larynx, and tongue, and then make a comprehensive overall impairment level judgment based on this. (4) Final response generation based on search enhancement: By designing search enhancement generation (RAG) technology, the assessment and suggestion generation process uses real-time and accurate access to professional knowledge such as Frenchay scale details and clinical literature to ensure that the output assessment conclusions and rehabilitation suggestions have authoritative basis, clinical rigor and traceable citation.
[0013] Other features and advantages of the invention will become clear from the following detailed description of exemplary embodiments of the invention with reference to the accompanying drawings. Attached Figure Description
[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.
[0015] Figure 1 This is a schematic diagram of the overall process of a speech disorder assessment method based on a multimodal large model according to an embodiment of the present invention; Figure 2 This is a flowchart of a speech disorder assessment method based on a multimodal large model according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the model reasoning and evaluation process based on the Frenchay scale according to an embodiment of the present invention; Figure 4 This is a schematic diagram of a large model knowledge enhancement method based on retrieval enhancement according to an embodiment of the present invention. Detailed Implementation
[0016] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the invention.
[0017] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.
[0018] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0019] In all the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0020] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0021] In summary, this invention aims to generate internal articulation motor representations consistent with articulation by completing the internal articulation organ missing modality, under the condition of only acquiring audio and lip video, reducing the dependence on direct acquisition by UTI (ultrasound tongue imaging) and improving scene adaptability; it constructs a multi-articulation organ assessment and reasoning system consistent with the Frinchay scale, enabling the large model to conduct itemized assessments by organ dimension and complete overall comprehensive judgment, outputting structured and verifiable assessment reports; it introduces RAG knowledge-enhanced retrieval technology, integrating Frinchay scale specifications, articulation disorder diagnosis and treatment guidelines, and literature evidence into the reasoning and suggestion generation process in real time, ensuring that the conclusions are professional, rigorous, interpretable, and citationable. This invention achieves intelligent assessment of articulation disorders based on the Frinchay scale through the collaborative processing of multimodal features, cross-modal mapping of the large model, and knowledge-enhanced reasoning.
[0022] See Figure 1 As shown, the overall process of the speech disorder assessment method based on a multimodal large model provided by this invention includes: First, feature encoding is performed on the input audio, lip video, and generated tongue ultrasound image. For example, the audio uses a self-supervised model (Hubert) to extract frame-level acoustic features to retain core information such as the energy of the articulation; the lip video uses a 3D CNN to extract the spatiotemporal features of lip shape and movement trajectory; the tongue ultrasound image is generated from the audio using a diffusion model, and similarly, the tongue spatial configuration features are extracted using a 3D CNN. Then, an MLP (Multilayer Perceptron) is used to align the features of the three types of features, ensuring that the multimodal features at the same time stamp correspond to the same articulation process. Next, the aligned features are input into a multimodal large model. For example, this model is based on the Qwen2.5-7B architecture and, after being fine-tuned under supervision using parallel datasets of audio, lip video, and tongue ultrasound, can establish a mapping relationship between external articulation performance and internal articulation movement. Then, the model calls the knowledge-enhanced retrieval module to retrieve professional assessment knowledge of the corresponding articulation organs in real time from a structured knowledge base containing guidelines for the diagnosis and treatment of articulation disorders and the Frenchay scale specifications. Finally, based on the Frenchay scale, the motor functions of organs such as the lips, jaw, palate, larynx, and tongue are graded and assessed, outputting the severity of articulation disorders and descriptions of specific organ problems. It should be noted that the large model in this paper can be a pre-trained language model with various architectures, such as a pre-trained language model based on the Transformer architecture.
[0023] Specifically, in combination Figure 1 and Figure 2 As shown, the provided speech disorder assessment method based on a multimodal large model includes the following steps: Step S1: Acquire multimodal data of the target, including audio signals, lip video, and tongue ultrasound images.
[0024] Audio signals contain phonetic parameters such as formant frequency, pitch, volume, aspiration time, voiced / voiced ratio, and phonetic length, which can reveal subtle pronunciation defects.
[0025] Lip video provides crucial visual motion information, enabling intuitive capture of mouth movement patterns during pronunciation. This overcomes the limitations of single-modal audio and enhances the comprehensiveness and accuracy of speech disorder assessment.
[0026] Ultrasound tongue imaging (UTI) provides real-time visualization of the sagittal dynamics of the tongue during continuous articulation, offering a crucial internal motion perspective for dysarthria assessment. UTI non-invasively captures internal tongue movements in real time, providing objective physiological evidence unavailable through traditional auditory processing, significantly improving the accuracy and interventional potential of speech disorder assessment. In this paper, UTI data was acquired from real-world data or generated using audio signals through a diffusion model.
[0027] Step S2 involves performing cross-modal feature extraction and semantic alignment on the audio signal, lip video, and tongue ultrasound image to map the extracted modal features to a unified semantic space aligned with the large model text embedding space, thereby obtaining a semantic vector sequence.
[0028] To address the mismatch between non-textual data such as audio and video and the understanding capabilities of large models, one embodiment designs a phased feature encoding and semantic alignment process.
[0029] First, targeted cross-modal feature extraction is performed. For example, for the input audio signal, a self-supervised learning-based speech model (such as HuBERT, Wav2Vec 2.0) is used to extract frame-level or phoneme-level deep acoustic features. These acoustic features can effectively capture key information related to the movement of the articulatory organs, such as the energy, pitch, and formants of the articulation. For the input lip video and ultrasound video, a 3D convolutional neural network (3D CNN) model is used to extract dynamic features that characterize the lips and ultrasound articulatory organs. This design enables the collaborative extraction of features from audio and lip video. The Whisper model is used to extract frame-level acoustic features from the audio, preserving the speaker's articulation characteristics and reducing noise interference. A model such as CLIP ViT-L / 14 is used to extract spatiotemporal features from the lip video, capturing the dynamic changes in lip movement and providing a basis for the generation of internal articulation movements. It should be noted that the 3D CNN used for feature extraction from lip video and ultrasound video can have the same or different structures, and this invention does not impose any restrictions on this.
[0030] Furthermore, the system performs feature mapping and alignment to the semantic space. Specifically, the extracted audio feature sequences and lip video feature sequences are mapped to a unified semantic space aligned with the large model's text embedding space through dedicated projection layers (such as multilayer perceptrons, MLPs). During this process, a large model language alignment loss function ensures alignment between audio events, corresponding actions in lip videos, and text within the semantic space, forming a temporally consistent multimodal "language." Additionally, when direct acquisition of ultrasound tongue image data is not possible, the system can integrate speech organ reverse reconstruction technology. Using a trained generative model (diffusion model), and conditioned by the aforementioned acoustic features, it generates corresponding, virtual tongue movement trajectory features in real time, and projects these features into the same unified semantic space. Through this design, audio, lip movements, and (real or generated) tongue movement information are all transformed into semantic vector sequences that the large model can directly "understand" and "process."
[0031] In summary, this invention designs a cross-modal semantic alignment and understanding technology. By utilizing a specific mapping network (such as a multilayer perceptron), the temporal features of non-textual modalities such as audio, lip video, and generated tongue ultrasound images are transformed into a unified semantic representation aligned with the text embedding space of a large language model. This enables the large model to understand and process multimodal information during the pronunciation process, achieving semantic-level fusion of heterogeneous data.
[0032] Step S3: Using the semantic vector sequence as input, a large model is used to simulate multi-level clinical reasoning to obtain preliminary assessment results of dysarthria. These preliminary assessment results include the severity level and type of dysarthria.
[0033] To enable the large model to simulate the structured diagnostic logic of clinicians, in one embodiment, a hierarchical, organ-oriented cueing engineering and reasoning framework was designed based on the Frenchay Dysarthria Assessment Scale. See also Figure 3 As shown, large models (such as Qwen, LLaMA, etc. after instruction fine-tuning) receive the above-aligned multimodal semantic sequence as input, and the evaluation process is broken down into two distinct stages.
[0034] In the first phase, a segmented assessment of the articulatory organs is performed. Specifically, a structured chain-of-thought prompt is injected into the large model, guiding it to independently analyze the functional status of the five core articulatory organs: lips, jaw, palate, larynx, and tongue. For example, for the "lips," the model combines semantic features from lip video to analyze their symmetry in a static state, and their range, strength, speed, and coordination during movement, outputting sub-conclusions consistent with clinical descriptions, such as "reduced lip range of motion, weak closure." This process can be performed in parallel or sequentially, generating a structured state description for each organ.
[0035] In the second phase, a comprehensive overall assessment is performed. Specifically, after completing all sub-assessments, the overall model is guided into the comprehensive reasoning phase. For example, by summarizing the sub-assessments of all organs and using the comprehensive scoring criteria of the Frenchay scale, the severity level of the overall dysarthria (e.g., normal, mild, moderate, severe) is inferred, and the main type of dysarthria (e.g., spastic, flaccid) is determined. This "sub-assessment first, then synthesis" reasoning path rigorously replicates the diagnostic thinking of clinicians, ensuring the logical consistency of the assessment process and the reliability of the results.
[0036] In summary, this invention mimics the reasoning process of speech therapists by designing a hierarchical clinical reasoning framework based on the Frenchay scale. It guides the large model to first conduct independent sub-assessments and problem descriptions of the five core articulatory organs—lips, jaws, palate, larynx, and tongue—before comprehensively determining the impairment level based on the sub-assessments, and finally outputting a structured assessment report, thus automating the clinical assessment process. This multi-organ assessment system based on the Frenchay scale better aligns with the reasoning rules of clinical assessment, achieving precise matching between multimodal features and corresponding dimensions of the scale, thereby realizing the transformation from organ status judgment to clinical problem description.
[0037] Step S4: Use the key information in the preliminary assessment results of the large model to set up a retrieval and query strategy to guide the large model to further generate the final assessment report and rehabilitation recommendations.
[0038] To ensure the professionalism, accuracy, and traceability of assessment conclusions and rehabilitation recommendations, this invention introduces a professional medical knowledge generation process based on retrieval enhancement technology, or a retrieval enhancement generation (RAG) module.
[0039] See Figure 4 As shown, the core of the retrieval enhancement generation module is a carefully constructed structured professional knowledge base, which indexes non-parametric knowledge such as the official manual of the Frenchay scale, clinical diagnosis and treatment guidelines for dysarthria, authoritative literature on rehabilitation training methods, and typical case reports.
[0040] In practical applications, once the large model completes its initial evaluation (whether it's a breakdown of individual conclusions or a comprehensive assessment), the RAG module is dynamically triggered. The specific process includes the following steps: Step S41, Query Construction: Transform key information from the current assessment conclusion of the large model (such as "lip closure weakness" and "moderate spastic dysarthria") into search queries.
[0041] Step S42, Hybrid Search: A hybrid strategy of "keyword exact matching + vector semantic similarity retrieval" is employed to retrieve the most relevant knowledge fragments from the professional knowledge base. Keyword retrieval ensures the accuracy of professional terminology, while vector retrieval enables the understanding of the semantics of clinical descriptions.
[0042] Step S43, Knowledge Fusion and Generation: The retrieved authoritative knowledge fragments (such as targeted lip muscle strength training methods for "weak lip closure") are used as additional context and input along with the original input and reasoning process of the large model to guide the model in generating the final assessment report and rehabilitation recommendations.
[0043] In one embodiment, the final output report can explicitly include: the level of impairment, a description of the specific problems in each organ, and the knowledge citation source corresponding to each rehabilitation suggestion (such as guideline name, chapter, document ID, etc.). This makes every output "verifiable," greatly enhancing clinical credibility and practicality, and also avoiding the model from generating "illusions" or unprofessional advice.
[0044] In summary, this invention constructs a multimodal fusion-based large-scale model for assessing articulation disorders, integrating audio, lip video, and generated tongue ultrasound image data. Through model optimization, it achieves deep fusion of information from internal and external speech organs, improving the accuracy, clinical suitability, and standardization of assessment results. Furthermore, a dynamic knowledge retrieval-enhanced interpretable generation method is designed. By integrating retrieval-enhanced generation technology, it can retrieve professional knowledge bases on articulation disorders (such as the Frenchay scale guidelines and clinical guidelines) in real time during the assessment process. The retrieved authoritative knowledge fragments are used as additional context to guide the model in generating assessment conclusions and rehabilitation suggestions with clearly cited sources, ensuring the professionalism and traceability of the output, and guaranteeing the citationability of the suggestions, thus ensuring clinical rigor.
[0045] Accordingly, the present invention also provides a speech disorder assessment system based on a multimodal large model to implement one or more aspects of the above-mentioned method. For example, the system includes: a multimodal data acquisition module for acquiring corresponding multimodal data for a target, wherein the multimodal data includes audio signals, lip videos, and tongue ultrasound images; a feature extraction and alignment module for performing cross-modal feature extraction and semantic alignment on the multimodal data to map the extracted modal features to a unified semantic space aligned with the large model's text embedding space, obtaining a semantic vector sequence; a preliminary assessment module for using the semantic vector sequence as input and simulating multi-level clinical reasoning logic using the large model to obtain a preliminary assessment result of the speech disorder, wherein the preliminary assessment result includes severity level and disorder type; and an assessment report generation module for setting a retrieval query strategy using key information in the preliminary assessment conclusion to guide the large model in generating an assessment report and rehabilitation recommendations. Each module in the system can be implemented using a general-purpose processor, a dedicated processor, an FPGA, or in combination with software.
[0046] In summary, compared with the prior art, the present invention has the following main advantages: (1) Existing technologies usually focus only on a single articulatory organ (such as the tongue) or a single modality analysis, resulting in incomplete assessment dimensions. This invention strictly follows the internationally accepted Frenchay Articulation Disorder Assessment Scale, constructs a complete assessment system covering the five core organs of the lips, jaw, palate, larynx, and tongue, and simulates the diagnostic logic of clinicians through hierarchical reasoning, making the assessment results more comprehensive, structured, and in line with clinical norms, and directly meeting the needs of clinical diagnosis and treatment.
[0047] (2) Existing technologies mainly rely on patterns in training data for judgment, lacking real-time professional knowledge support, which can easily lead to inaccurate or non-clinical recommendations. This invention integrates RAG technology to call authoritative professional knowledge bases in real time during the reasoning process, ensuring that every assessment conclusion and rehabilitation recommendation has clear clinical basis and reference source, which greatly improves the professionalism, rigor and clinical acceptability of the assessment results.
[0048] (3) Existing technologies often simply splice or analyze multimodal data, lacking an effective cross-modal alignment mechanism. This invention achieves precise alignment and deep fusion of audio, external lip movements, and internal vocal organ movements in both the temporal and semantic dimensions by designing a specialized temporal feature extraction and semantic space mapping network. Through the deep fusion and precise alignment of multimodal information, a more reliable multimodal evidence foundation is provided for accurate evaluation.
[0049] (4) Compared with traditional manual assessment or existing semi-automatic systems, this invention relies on the rapid reasoning capability of multimodal large models to process input data in real time and complete the entire process from data input to detailed assessment report generation within minutes, which greatly improves assessment efficiency and provides technical feasibility for large-scale screening and frequent rehabilitation progress tracking.
[0050] (5) This invention enables an end-to-end assessment system that supports modal missing data, achieving complete automated assessment. It maintains input audio, lip video, and tongue ultrasound at all times. Furthermore, in the absence of tongue ultrasound images, tongue movement information can be supplemented through ultrasound reverse reconstruction, achieving fully automated processing from multimodal data input to structured assessment report generation. Verification has shown that this invention improves the accuracy, real-time performance, and robustness of speech disorder assessment.
[0051] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.
[0052] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0053] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0054] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0055] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0056] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It will be known to those skilled in the art that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are equivalent.
[0057] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the invention is defined by the appended claims.
Claims
1. A speech disorder assessment method based on a multimodal large model, comprising the following steps: For the target, acquire the corresponding multimodal data, which includes audio signals, lip video, and tongue ultrasound images; Cross-modal feature extraction and semantic alignment are performed on the multimodal data to map the extracted modal features to a unified semantic space aligned with the large model text embedding space, thereby obtaining a semantic vector sequence. Using the semantic vector sequence as input, the large model is used to simulate multi-level clinical reasoning logic to obtain preliminary assessment results of articulation disorders, including severity level and disorder type. Using the key information from the preliminary assessment conclusions, a retrieval strategy is set up to guide the large model in generating assessment reports and rehabilitation recommendations.
2. The method according to claim 1, characterized in that, Cross-modal feature extraction and semantic alignment of the multimodal data includes: The audio signal, lip video and tongue ultrasound image were respectively coded to obtain the corresponding acoustic features, spatiotemporal features reflecting lip shape and movement trajectory, and tongue spatial configuration features. The acoustic features, the spatiotemporal features reflecting lip shape and movement trajectory, and the tongue spatial configuration features are aligned using a multilayer perceptron so that each modal feature at the same timestamp corresponds to the same pronunciation process.
3. The method according to claim 2, characterized in that, The acoustic features are deep acoustic features at the frame or phoneme level extracted using a self-supervised model. The spatiotemporal features reflecting lip shape and motion trajectory are extracted using a first 3D convolutional neural network, and the tongue spatial configuration features are extracted using a second 3D convolutional neural network.
4. The method according to claim 1, characterized in that, The tongue ultrasound image is either real ultrasound tongue image data collected, or tongue motion trajectory features generated using a trained generative model based on acoustic features extracted from the audio signal.
5. The method according to claim 1, characterized in that, The preliminary assessment results were obtained according to the following steps: The large model is injected with structured thought chain prompts to guide it to independently analyze the functional state of multiple speech organs in turn, generate a structured state description for each speech organ, and obtain the sub-evaluation conclusions for each speech organ. The large model summarizes the sub-assessment conclusions, and based on the comprehensive scoring criteria of the Frenchay scale, infers the severity level of the overall articulation disorder and determines the disorder type.
6. The method according to claim 1, characterized in that, The assessment report and rehabilitation recommendations were obtained through the following steps: Transform the key information in the current evaluation conclusion of the large model into a retrieval query; A hybrid strategy combining keyword matching and vector semantic similarity retrieval is adopted to retrieve the most relevant knowledge fragments from the professional knowledge base; The retrieved knowledge fragments are used as additional context, and are input along with the original input and reasoning process of the large model to guide the large model in generating the final assessment report and rehabilitation recommendations. The assessment report and rehabilitation recommendations include the level of impairment, the problem description of each organ, and the knowledge reference source corresponding to each rehabilitation recommendation.
7. The method according to claim 1, characterized in that, The large model is a pre-trained language model based on the Transformer architecture. After being fine-tuned under the supervision of parallel datasets of audio, lip video, and tongue ultrasound, a mapping relationship between external articulation performance and internal articulation movement is established.
8. The method according to claim 5, characterized in that, The multiple vocal organs include the lips, jaw, palate, larynx, and tongue.
9. A speech disorder assessment system based on a multimodal large model, comprising: Multimodal data acquisition module: used to acquire corresponding multimodal data for a target, including audio signals, lip video and tongue ultrasound images; Feature extraction and alignment module: used to perform cross-modal feature extraction and semantic alignment on the multimodal data, so as to map the extracted modal features to a unified semantic space aligned with the large model text embedding space, and obtain a semantic vector sequence; Preliminary assessment module: This module uses the semantic vector sequence as input and the large model to simulate multi-level clinical reasoning logic to obtain preliminary assessment results of articulation disorders. The preliminary assessment results include severity level and disorder type. Assessment report generation module: This module is used to set up a retrieval and query strategy using key information from the preliminary assessment conclusions, and to guide the large model to generate an assessment report and rehabilitation recommendations.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Speech and language hypoacousie multi-parameter diagnosis and rehabilitation apparatus and cloud rehabilitation system
CN105982641A