Sound structuring abnormity positioning method and device, electronic equipment and storage medium

By using a multimodal attention mechanism and a large language model for cross-modal fusion between audio and video, the accuracy and efficiency issues in the assessment of people with articulation disorders in existing technologies have been resolved, enabling precise localization of word-level articulation abnormalities and support for personalized rehabilitation programs.

CN121528255AActive Publication Date: 2026-02-13PEKING UNION MEDICAL COLLEGE HOSPITAL +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202610056268.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-16
Publication Date
2026-02-13
Estimated Expiration
2046-01-16

AI Technical Summary

Technical Problem

Existing technologies rely on subjective evaluation when assessing speech rehabilitation for people with articulation disorders, which consumes a lot of human resources and makes it difficult to achieve unified assessment standards. Furthermore, single-modal recognition methods are difficult to accurately identify and locate word-level articulation errors.

Method used

A multimodal attention mechanism is used to establish a mapping between audio and lip video. Through cross-modal fusion, a multimodal large language model is used to locate word-level articulation anomalies, and audio and video features are combined for refined judgment.

Benefits of technology

It enables precise localization of word-level articulation abnormalities, improves the accuracy and efficiency of assessment, provides clear rehabilitation training targets, reduces the risk of missed detection and misdiagnosis, and supports the development of personalized rehabilitation plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121528255A_ABST
    Figure CN121528255A_ABST
Patent Text Reader

Abstract

The invention provides a sound structuring abnormity positioning method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring audio and video data recorded when a target object speaks; performing text conversion and correction processing on audio data in the audio and video data to obtain text data; performing lip region identification and extraction on video data in the audio and video data to obtain lip video data; extracting an audio segment and a video segment corresponding to each word in the text data from the audio data and the lip video data to obtain to-be-positioned data; and inputting the to-be-positioned data into an audio encoder and a video encoder included in the multi-modal large language model subjected to fine adjustment of the target training data, and performing cross-modal fusion through a unified embedding space to obtain a phonetic composition marking result. According to the method and the device, a single-mode recognition blind area is made up, the accuracy and the reliability of pronunciation state judgment are improved, and pronunciation abnormity of each character can be accurately positioned.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of multi-modal information processing, in particular to a speech abnormality positioning method and device, electronic equipment and storage medium. BACKGROUND

[0002] Currently, when evaluating the speech rehabilitation of people with speech disorders (such as patients with cleft lip and palate), it mainly relies on experienced speech therapists to listen and label. This subjective evaluation method not only consumes a large amount of human resources, but also makes it difficult to achieve uniform evaluation standards. In addition, the evaluation of nasalization and other features also needs to be completed by clinically trained personnel. In order to improve the evaluation efficiency, researchers have developed a technology that uses acoustic features and machine learning algorithms to automatically detect speech errors. However, this technology still cannot accurately identify and precisely locate word-level speech errors. SUMMARY

[0003] The embodiments of the present application provide a speech abnormality positioning method and device, electronic equipment and storage medium to solve the problem of how to position word-level speech abnormalities for speech abnormality patients. The present application first proposes a multi-modal unified embedding structure for Chinese word-level speech positioning, which uses a cross-modal attention mechanism to establish a mapping between audio, lip video and text, thereby achieving fine-grained judgment of speech abnormality categories and their time positions.

[0004] In a first aspect, the embodiments of the present application provide a speech abnormality positioning method, which comprises: obtaining audio and video data recorded when a target object speaks; performing text conversion and correction processing on the audio data in the audio and video data to obtain text data; performing lip region recognition and extraction on the video data in the audio and video data to obtain lip video data; extracting an audio segment and a video segment corresponding to each word in the text data from the audio data and the lip video data to obtain to-be-positioned data; inputting the to-be-positioned data into an audio encoder and a video encoder included in a multi-modal large language model fine-tuned by target training data, performing cross-modal fusion through a unified embedding space, and obtaining a speech marking result, wherein the speech marking result includes a speech state corresponding to each word in the text data; the target training data includes a plurality of sample audio and video data labeled with word-level speech states, and the recording object of the sample audio and video data and the target object belong to the same speech disorder type.

[0005] In a second aspect, the embodiments of the present application also provide a speech abnormality positioning device, which comprises: The first obtaining module is configured to obtain audio and video data recorded when the target object speaks. The first processing module is configured to perform text conversion and correction processing on audio data in the audio and video data to obtain text data. The second processing module is configured to perform lip region recognition and extraction on video data in the audio and video data to obtain lip video data. The extraction module is configured to extract, from the audio data and the lip video data, an audio segment and a video segment corresponding to each word in the text data to obtain to-be-positioned data. The speech construction positioning module is configured to input the to-be-positioned data into an audio encoder and a video encoder included in a multi-modal large language model fine-tuned by target training data, perform cross-modal fusion through a unified embedding space, and obtain speech construction marking results, wherein the speech construction marking results include speech construction states corresponding to each word in the text data.

[0006] In a third aspect, an electronic device is provided, which includes a processor, a memory, and a computer program stored in the memory and executable on the processor, and the computer program, when executed by the processor, implements the speech construction abnormality positioning method described above.

[0007] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program, when executed by a processor, implements the speech construction abnormality positioning method described above.

[0008] The embodiments of the present application at least have the following technical effects: The technical scheme of the embodiments of the present application collects audio and video data when a target object speaks, performs audio text conversion and correction and video lip region extraction, splits out audio and video segments corresponding to each word in the text, and then processes them by using a multi-modal large language model fine-tuned by samples of the same speech construction disorder type, to realize word-level speech construction state marking. The present application compensates for the blind area of single-modal recognition by fusing acoustic and lip visual features through multi-modal fusion, improves the accuracy and reliability of speech construction state determination, accurately positions the speech construction abnormality of each word, and provides a clear target for rehabilitation training. Meanwhile, the model is fine-tuned to adapt to a specific speech construction disorder type, which strengthens the sensitivity to pathological pronunciation patterns, reduces the risk of missed detection and misjudgment of a general model, and outputs structured speech construction marking results, which intuitively presents abnormal words and corresponding states, efficiently supports clinical evaluation and individualized rehabilitation program development, and greatly improves the accuracy, pertinence, and clinical conversion efficiency of speech construction abnormality positioning. BRIEF DESCRIPTION OF DRAWINGS

[0009] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0010] Figure 1 This is a flowchart illustrating the articulation anomaly localization method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the articulation abnormality localization device provided in the embodiments of this application; Figure 3 A block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0011] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0012] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0013] In the various embodiments of this application, it should be understood that the sequence number of each process described below does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0014] In related technologies, speech rehabilitation assessments for individuals with articulation disorders (such as those with cleft lip and palate) primarily rely on experienced speech therapists for listening and annotation. This subjective evaluation method not only consumes significant human resources but also struggles to achieve standardized assessment criteria. Furthermore, assessing features such as nasalization requires rigorously trained clinicians. To improve assessment efficiency, researchers have developed a technique that utilizes acoustic features and machine learning algorithms to automatically detect articulation errors.

[0015] While existing automated methods can detect the degree of nasalization or classify speech as normal or abnormal, most rely on pure audio features, neglecting the important information of the speaker's mouth shape. Due to their structural defects, patients with cleft lip and palate often exhibit unique lip movements during pronunciation, making it difficult for simple acoustic analysis to distinguish similar phonemes or pinpoint specific errors. Furthermore, existing pronunciation error detection systems often target isolated syllables or short words, lacking the ability to annotate each word in natural conversation.

[0016] like Figure 1 As shown in the embodiment of this application, a method for locating articulation abnormalities is provided, the method comprising: Step 101: Obtain the audio and video data recorded when the target person is speaking.

[0017] This invention provides a method for locating articulation abnormalities, applied to an articulation abnormality location system. This system can be integrated into mobile terminals, remote consultation platforms, or existing hospital information systems. The system can acquire audio and video data for articulation abnormality location, specifically files recorded during the target's speech. Acquisition can be performed after recording or in real-time during recording. Specifically, the audio and video data acquired by the system can also be recorded by the user via mobile phone and transmitted to the system via the cloud, addressing the inconvenience of offline articulation abnormality location.

[0018] Specifically, while the target subject is speaking, facial video and voice signals are captured using a camera and microphone simultaneously, or the same recording device, to obtain audio and video data. To reduce distance attenuation and environmental noise, the audio sampling frequency should be higher than 16kHz; to ensure that the details of lip movements are not blurred and the timing information is complete, the video sampling frequency should be higher than 30FPS. For the recording scenario, to avoid noise interference and the failure of lip area recognition due to backlighting and shadows, the recording process must be carried out in a quiet, echo-free indoor environment, thereby effectively reducing interference factors in the acquired audio and video data, reducing the complexity of subsequent data processing, and improving system robustness.

[0019] The spoken content of the target subject can be standardized text to ensure coverage of various articulation scenarios. It can also include environments such as free speech and sentence repetition to capture natural articulation performance. Standardized recording content ensures the comparability of recording data from different subjects and at different times, providing a unified benchmark for anomaly localization.

[0020] It should be noted that for children with speech disorders, interesting texts can be designed to reduce the resistance to recording. The interesting texts can be nursery rhymes or picture book contents. For patients with severe speech disorders, the recording content can be simplified and the recording time can be extended to obtain sufficient samples.

[0021] Step 102: Perform text conversion and correction processing on the audio data in the audio-visual data to obtain text data.

[0022] After obtaining the audio-visual data, for the audio data in the audio-visual data, speech recognition technology is used, such as auditory large language models like Qwen2.5Omni or SALMONN-video, and data augmentation is performed on cleft lip and palate speech to improve recognition fairness. The pronunciation audio of the target object is converted into the original text. Considering that the pronunciation of patients with speech disorders may be blurred or misaligned, a model that supports dialect adaptation or optimization of abnormal speech sounds needs to be selected, such as an end-to-end speech recognition model based on deep learning. By pre-training a large amount of normal speech and abnormal speech data, the recognition ability for non-standard pronunciation is improved. Further, the errors in the original text are corrected to ensure that the text data is consistent with what the target object actually wants to express.

[0023] Specifically, when performing correction processing, errors with unclear semantics and similar pronunciations can be automatically corrected based on the language model, such as correcting the recognized "three peaks" to "mountain peak". For ambiguous texts that cannot be solved by automatic correction, annotators can combine audio playback and clinical experience for correction.

[0024] Accurate text data provides a basis for subsequent alignment processing, ensuring that each character can find the corresponding acoustic and visual features, and avoiding positioning deviations caused by text errors.

[0025] It should be noted that before performing text conversion and correction processing on the audio data, preprocessing of the audio data can be performed, including noise suppression, silent segment detection, and speech endpoint detection.

[0026] Step 103: Perform lip region recognition and extraction on the video data in the audio-visual data to obtain lip video data.

[0027] For the video data in the audio-visual data, first, a face detection algorithm is used to locate the face region of the target object in each frame of the video, determine the bounding box coordinates of the face, and define the range for lip region extraction. Then, based on face key point detection technology, facial key points are recognized, key points related to the lips are selected, and the polygonal boundary of the lip region is fitted through these key points. Finally, the region is extracted through image cropping technology to obtain lip video data that only contains the lips.

[0028] By accurately separating the lip region from the complete video, interference from other parts of the face and the background environment can be excluded, improving the accuracy of subsequent articulation state judgment.

[0029] Step 104: Extract the audio segment and video segment corresponding to each character in the text data from the audio data and the lip video data to obtain the data to be located.

[0030] After obtaining the text data and the lip video data, the text data is segmented. For Chinese, it is mainly segmented into single characters. If it is a polysyllabic character, it is split into syllables. Then, based on the timestamp information output by the speech recognition model, the start time and end time of each character in the audio data are determined. For example, the audio segment corresponding to the text "爸" is 0.5 - 0.8 seconds. According to the timestamp of each character, the corresponding audio segment is cropped from the original audio data to ensure that each audio segment contains only the pronunciation of a single character. Based on the audio-video synchronization, the audio timestamp of each character is mapped to the video data to determine the video frame range corresponding to the pronunciation of the character. For example, the video frames corresponding to 0.5 - 0.8 seconds are frames 15 - 24. Then, this frame range is extracted from the lip video data to obtain the video segment synchronized with the audio segment. Finally, the audio segment and video segment corresponding to each character in the text data are determined as the data to be located.

[0031] Step 105: Input the data to be located into the audio encoder and video encoder included in the multi-modal large language model fine-tuned by the target training data, and perform cross-modal fusion through a unified embedding space to obtain an articulation marking result, where the articulation marking result includes the articulation state corresponding to each character in the text data; the target training data includes multiple sample audio-video data marked with the articulation state at the character level, and the recording object of the sample audio-video data and the target object belong to the same type of articulation disorder.

[0032] In the embodiment of the present application, a large language model with three-modal input of audio, video, and text is selected, such as a multi-modal model based on the Transformer architecture. This type of model has powerful cross-modal feature fusion capabilities and semantic understanding capabilities, and can process acoustic features, visual features, and text features simultaneously. In the embodiment of the present application, target training data is also pre-determined. The target training data includes a large number of sample audio-video data marked with the articulation state at the character level, and the recording object of the sample audio-video data and the target object belong to the same type of articulation disorder. The annotation content needs to be completed by a professional speech therapist, marking the articulation state of each character. The articulation state includes normal and abnormal, and the abnormal also includes the corresponding abnormal type, such as substituting a stop sound in the throat, excessive nasalization, deviation of alveolar fricatives, etc. For different types of articulation disorders, the types of articulation abnormalities are different.

[0033] In this embodiment, the model input to the data to be localized is a multimodal large language model fine-tuned with target training data. After obtaining the data to be localized, it is input into the multimodal large language model fine-tuned with target training data. The multimodal large language model includes an audio encoder and a video encoder. The multimodal large language model can perform cross-modal fusion of the audio data input to the audio encoder and the video data input to the video encoder through a unified embedding space to obtain articulation marking results. The articulation marking results include the articulation state corresponding to each word in the text data. The articulation state includes normal and abnormal. When the articulation state is abnormal, it further includes the articulation abnormality type.

[0034] This application, through the acquisition of audio and video data of the target subject speaking, performs audio-to-text conversion correction and video lip region extraction to extract the audio and video segments corresponding to each character of the text. Then, using a multimodal large language model fine-tuned with samples of isomorphic articulation disorders, word-level articulation status marking is achieved. This application, by fusing acoustic and lip visual features through multimodal integration, compensates for the blind spots of single-modal recognition, improving the accuracy and reliability of articulation status determination. It can accurately locate the articulation abnormality of each character, providing a clear target for rehabilitation training. Simultaneously, the targeted fine-tuning of the model adapts to specific articulation disorder types, enhancing sensitivity to pathological pronunciation patterns and reducing the risk of missed detections and misjudgments by general models. The output structured articulation marking results intuitively present abnormal characters and their corresponding states, efficiently supporting clinical assessment and personalized rehabilitation plan development, significantly improving the accuracy, targeting, and clinical translation efficiency of articulation abnormality localization.

[0035] It should be noted that the articulation abnormality localization method provided in this application is applicable to speech rehabilitation assessment of patients with cleft lip and palate, patients with cerebral palsy and speech disorders, etc., and is also applicable to other speech pathology scenarios that require precise annotation of pronunciation errors.

[0036] In an optional embodiment of this application, extracting the audio segment and video segment corresponding to each character in the text data from the audio data and the lip video data includes: The audio data and the lip video data are time-aligned; The audio data and the text data are aligned at the character level to determine the audio segment corresponding to each character in the text data, as well as the start and end times of the audio segment. For each character in the text data, the corresponding video segment is extracted from the lip video data based on the start and end times of the character.

[0037] Time alignment is performed on audio data and lip video data to eliminate timing discrepancies between audio and video during recording or transmission, ensuring that the acoustic and visual signals corresponding to the same vocalization action are completely synchronized on the time axis, and ensuring that the extracted audio and video segments reflect the characteristics of the same vocalization action.

[0038] The audio and text data are aligned at the word level to establish a time mapping relationship between text word units and audio signals. The continuous audio stream is decomposed into discrete segments based on words to obtain the audio segment corresponding to each word in the text data, as well as the start and end times of the audio segment.

[0039] Based on a defined character-level timestamp, video segments corresponding to the time period are extracted from the aligned lip video data to ensure that the video segments can accurately reflect the lip movement characteristics when a single character is pronounced.

[0040] The above-described implementation scheme of this application achieves precise association between multimodal data and text character units through time alignment, character-level binding, and video extraction, and decomposes continuous audio and video streams into discrete feature units corresponding to each Chinese character. This ensures that the acoustic features of each character's pronunciation are completely matched with the visual features of lip movement without confusion, providing a pure and reliable feature carrier for subsequent precise analysis of character-level phonetic states.

[0041] In an optional embodiment of this application, the data to be located is input into the audio encoder and video encoder of a multimodal large language model fine-tuned by the target training data, and cross-modal fusion is performed through a unified embedding space to obtain the articulation tagging result, including: Feature extraction is performed on the audio and video segments corresponding to each character in the data to be located, to obtain the audio feature vector and visual feature vector corresponding to each character; The audio feature vector and visual feature vector corresponding to each character are concatenated with the modal marker, and the concatenation result is input into the multimodal large language model. Obtain the pronunciation type embedding vector corresponding to each character output by the multimodal large language model; Based on the pronunciation type embedding vector corresponding to each character, and using the articulation state classification module, the articulation state corresponding to each character is determined, and the articulation marking result is obtained.

[0042] First, the data to be located is transformed into a vector form that can be processed by a multimodal large language model. For each character's audio segment, acoustic features such as Mel spectrum, formants, and perturbations are extracted to obtain an audio feature vector. For each character's video segment, visual features are extracted using a convolutional neural network to obtain a visual feature vector.

[0043] Feature extraction for audio segments includes extracting: Mel-frequency cepstral coefficients, reflecting the spectral envelope characteristics of pronunciation; fundamental frequency, characterizing the vocal cord vibration frequency; and audio energy and spectral entropy, reflecting the intensity and spectral complexity of pronunciation. Simultaneously, since articulation is a dynamic process, frame-level acoustic features need to be temporally aggregated. Through methods such as sliding windowing, temporal averaging, or recursive feature fusion, features from consecutive frames are integrated into a global audio feature vector for a single word, ensuring that the features contain complete temporal information about pronunciation.

[0044] Feature extraction from video segments includes lip movement geometry, visual texture, and morphological features. Based on the keypoint coordinates of the lip region, geometric parameters such as lip opening and closing, corner stretching, and lip area change rate are calculated. Motion trajectory features are extracted through the dynamic changes of the keypoint coordinate sequence. Image feature extraction algorithms are used to extract texture features of the lip region and morphological features of the lips. For video segments, dynamic features of lip movement are extracted through inter-frame difference analysis, reflecting the smoothness and coordination of lip movements during articulation. The extracted visual features are normalized to eliminate feature bias caused by individual differences, ensuring that the model focuses on the differences in articulation movements themselves.

[0045] Furthermore, to achieve effective differentiation of multimodal features and model adaptation, modal tags are used to clarify the modal attributes of features, thereby activating the visual and audio encoders, avoiding cross-modal feature confusion, and providing a standardized input format for multimodal large language models. Modal tags are identifiers used to distinguish audio and visual features, and must be concise and unique. Specifically, modal tags can be fixed-dimensional trainable vectors with the same dimensions as the feature vectors, and are pre-concatenated before the feature vectors to form the final input sequence.

[0046] The audio feature vector of each word is concatenated with the audio modality marker element by element, and the visual feature vector is concatenated with the visual modality marker element by element. The two fused vectors are combined in a fixed order to form the final input of the model, ensuring that the input format is compatible with the multimodal input format during model pre-training.

[0047] Multimodal large language models transform feature-tag fusion vectors into pronunciation type embedding vectors with strong discriminative power.

[0048] Finally, the semantically discriminative embedding vectors are mapped to explicit articulation states through the articulation state classification module. This module needs to adapt to the feature distribution of the embedding vectors and be optimized to meet the classification requirements of articulation states.

[0049] The above-described implementation scheme of this application extracts character-level audio and visual core features, combines them with modal tagging to achieve ordered input of multimodal information, performs deep fusion and embedding mapping using a multimodal large language model, and then completes the articulation state determination through a classification module. Modal tagging guides the model to accurately identify different modal information. After fine-tuning the pre-trained model with data of the same type of articulation disorder, the generated pronunciation type embedding vector can accurately represent pathological pronunciation patterns. Combined with a dedicated classification module, this significantly improves the accuracy and specificity of character-level articulation state determination. Ultimately, it achieves fine-grained localization of articulation abnormalities, clearly distinguishes normal pronunciation from various abnormal types, provides objective data support for clinical diagnosis, reduces the misjudgment rate of model inference, and improves the pertinence and scientific nature of rehabilitation training guidance.

[0050] In an optional embodiment of this application, the audio data in the audio and video data undergoes text conversion and correction processing to obtain text data, including: Obtain the script containing the content of the target object's speech; The audio data is preprocessed and then converted into text to obtain the original text data; The original text data is matched and corrected based on the content script to obtain the text data.

[0051] Specifically, when determining the text data, it is necessary to obtain the content script. The content script can be a standardized script provided by the staff before recording. The target audience reads the script aloud to ensure that the pronunciation content is completely consistent with the script.

[0052] Then, through audio preprocessing, interference is eliminated and pronunciation features are enhanced to provide high-quality audio signals for text conversion. The preprocessed audio data is then converted into text to obtain the original text data.

[0053] Finally, by utilizing the deterministic semantics and sentence structure of the script, targeted corrections are made to the transcription errors in the original text, ultimately outputting standardized text that is highly consistent with the actual pronunciation.

[0054] The above-described implementation scheme of this application optimizes signal quality through audio preprocessing, generates original text by adapting transcription techniques to address articulation anomalies, and then uses the deterministic semantics and sentence structure of the script to complete accurate matching and correction, forming text data that provides a reliable text benchmark for subsequent word-level audio and video segment alignment and articulation state analysis.

[0055] In an optional embodiment of this application, lip region recognition and extraction are performed on the video data in the audio and video data to obtain lip video data, including: Perform facial recognition on the video data to obtain facial video data; The key points of the lip region are located and extracted from the face video data to obtain the lip video data.

[0056] First, a face detection algorithm is used to locate the face region of the target object in each frame of the video, determine the bounding box coordinates of the face, and obtain face video data. Then, based on facial landmark detection technology, facial landmarks are identified, and landmarks related to the lips are selected. The polygonal boundary of the lip region is fitted by these landmarks, and the region is extracted by image cropping technology to obtain lip video data containing only the lips.

[0057] The above-described implementation scheme of this application uses the face as a stable reference system, effectively avoiding interference from irrelevant information such as environmental background and other facial organs. At the same time, through temporal optimization and adaptation to special scenes, it ensures the accuracy and stability of the lip region positioning, so that the extracted video data can completely retain key visual features such as lip shape changes and movement trajectories during the articulation process, laying a reliable visual foundation for the accurate analysis of the articulation state.

[0058] In an optional embodiment of this application, after obtaining the articulation marking results, the method further includes: Based on the articulation marking results, obtain the erroneous words with abnormal articulation states, the articulation abnormality type corresponding to the erroneous words, and their occurrence positions in the audio and video data; Obtain correction suggestions that match the articulation anomaly type of the erroneous word; Based on the type of articulation error, location of occurrence, and correction suggestions corresponding to the erroneous word, an articulation error location report is generated.

[0059] In this embodiment of the application, after obtaining the phonetic marking results, the phonetic state of each character is traversed, erroneous characters with abnormal phonetic states are filtered out, and the phonetic abnormality type corresponding to each erroneous character and the occurrence position of the erroneous character in the audio and video data are obtained. The occurrence position includes the start time and end time in the audio data, and the start frame and end frame in the video data.

[0060] Furthermore, based on the articulation abnormality type of the erroneous word, corresponding correction suggestions are extracted from a pre-built, standardized, and structured rule-based correction suggestion library. In this library, different articulation abnormality types correspond to different correction suggestions. One articulation abnormality type can correspond to multiple different correction suggestions. Matching correction suggestions can be determined based on the individual characteristics of the target individual, ensuring the suggestions are targeted and improving the efficiency of rehabilitation training.

[0061] Finally, the articulation anomaly type, location, and correction suggestions corresponding to each erroneous word in the articulation marking results are integrated into a structured, visual, and easy-to-understand articulation anomaly localization report.

[0062] Specifically, the articulation disorder localization report may include a summary, specifically the total number of erroneous words, the main types of articulation disorders, and the key directions for rehabilitation training, to quickly understand the core information of this test. The report also includes basic information, specifically the target subject's basic information and recording information. Furthermore, the report includes a detailed analysis of articulation disorders, specifically the type of disorder, location of occurrence, and description of abnormal characteristics for each erroneous word. Finally, the report includes personalized correction suggestions, specifically pronunciation guidance, training methods, assistive tools, and training plans.

[0063] The above-mentioned implementation scheme of this application is based on the results of articulation marking. By extracting erroneous words, abnormality types and their locations, it establishes a mapping between abnormal information and personalized correction suggestions. Finally, it integrates and generates a structured and easy-to-understand articulation abnormality localization report, which achieves a clear presentation of articulation abnormalities and provides intervention guidance through targeted correction suggestions.

[0064] In an optional embodiment of this application, after obtaining the articulation marking results, the method further includes: On the articulation error location display interface, the error words with an abnormal articulation status are displayed; Upon receiving a playback request for the erroneous word, the video and audio segments corresponding to the erroneous word are played.

[0065] The articulation anomaly localization system of this application provides an articulation anomaly localization display interface. After generating an articulation anomaly report, the interface displays erroneous words with an abnormal articulation status. For example, the entire text of the text data can be displayed on the articulation anomaly localization display interface, highlighting erroneous words to visually distinguish them from normal words. Highly recognizable marking methods are used, such as red highlighting, underlining, bolding special fonts, and icon embellishments, while avoiding excessive marking that would clutter the interface.

[0066] Specifically, the system displays associated anomaly information near the erroneous word, such as the type of articulation error, its location, and severity. Furthermore, the articulation error location display interface provides a playback control. Users can select the erroneous word to be played back and click the playback control to trigger a playback request. Upon receiving the playback request, the articulation error location system plays the corresponding video and audio segments, visually reproducing the acoustic and lip movement characteristics of the incorrect pronunciation, providing a reference for anomaly analysis and correction. Simultaneously, after playing the video and audio segments corresponding to the erroneous word, the system can play the corresponding standard pronunciation audio and video, facilitating user practice and assisting in self-correction.

[0067] The above-described implementation scheme of this application transforms text information with abnormal pronunciation into an intuitive and perceptible interactive experience through visual interaction and playback of erroneous words, thereby improving the user experience.

[0068] The articulation anomaly localization method provided by the embodiments of this application has been described above. The articulation anomaly localization device provided by the embodiments of this application will be described below with reference to the accompanying drawings.

[0069] like Figure 2 As shown, this embodiment of the invention also provides a device for locating articulation abnormalities, the device comprising: The first acquisition module 210 is used to acquire audio and video data recorded when the target object is speaking; The first processing module 220 is used to perform text conversion and correction processing on the audio data in the audio and video data to obtain text data; The second processing module 230 is used to perform lip region recognition and extraction on the video data in the audio and video data to obtain lip video data. Extraction module 240 is used to extract the audio segment and video segment corresponding to each character in the text data from the audio data and the lip video data to obtain the data to be located; The articulation localization module 250 is used to input the data to be localized into the audio encoder and video encoder included in the multimodal large language model fine-tuned by the target training data, and to perform cross-modal fusion through a unified embedding space to obtain articulation marking results. The articulation marking results include the articulation state corresponding to each character in the text data. The target training data includes multiple sample audio and video data labeled with character-level articulation states. The recording objects of the sample audio and video data and the target object belong to the same articulation disorder type.

[0070] Optionally, the extraction module includes: The first processing submodule is used to perform time alignment on the audio data and the lip video data; The first determining submodule is used to perform character-level alignment on the audio data and the text data, and to determine the audio segment corresponding to each character in the text data, as well as the start and end times of the audio segment. The second determining submodule is used to extract the video segment corresponding to each character in the text data from the lip video data based on the start and end times corresponding to the character.

[0071] Optionally, the articulation localization module includes: The feature extraction submodule is used to extract features from the audio and video segments corresponding to each character in the data to be located, and to obtain the audio feature vector and visual feature vector corresponding to each character. The input submodule is used to concatenate the audio feature vector and visual feature vector corresponding to each character with the modal marker, and input the concatenation result into the multimodal large language model; The first acquisition submodule is used to acquire the pronunciation type embedding vector corresponding to each character output by the multimodal large language model; The third determining submodule is used to determine the articulation state of each character based on the embedded vector of the pronunciation type corresponding to each character and the articulation state classification module, so as to obtain the articulation marking result.

[0072] Optionally, the first processing module includes: The second acquisition submodule is used to acquire the content script when the target object speaks; The conversion submodule is used to preprocess the audio data and perform text conversion to obtain the original text data; The correction submodule is used to perform matching and correction on the original text data based on the content script to obtain the text data.

[0073] Optionally, the second processing module includes: The recognition submodule is used to perform face recognition on the video data to obtain face video data; The region extraction submodule is used to locate and extract key points in the lip region of the face video data to obtain the lip video data.

[0074] Optionally, after obtaining the articulation marking results, the device further includes: The determination module is used to determine, based on the articulation marking results, the erroneous words with abnormal articulation states, the articulation abnormality type corresponding to the erroneous words, and their occurrence positions in the audio and video data; The second acquisition module is used to acquire correction suggestions that match the articulation abnormality type of the erroneous word; The generation module is used to generate a pronunciation abnormality location report based on the pronunciation abnormality type, occurrence location, and correction suggestions corresponding to the erroneous word.

[0075] Optionally, after obtaining the articulation marking results, the device further includes: The display module is used to display the erroneous words with abnormal articulation status on the articulation abnormality location display interface; The playback module is used to play the video and audio segments corresponding to the erroneous word when a playback request for the erroneous word is received.

[0076] The articulation disorder localization device provided in this application collects audio and video data of the target subject speaking, performs audio-to-text conversion and correction, and extracts the lip region from the video, splitting the audio and video segments corresponding to each character of the text. Then, it uses a multimodal large language model fine-tuned with samples of the same articulation disorder type to infer and achieve character-level articulation status marking. This application compensates for the blind spots of single-modal recognition by fusing acoustic and lip visual features through multimodal integration, improving the accuracy and reliability of articulation status judgment. It can accurately locate the articulation disorder of each character, providing a clear target for rehabilitation training. At the same time, the targeted fine-tuning of the model is adapted to specific articulation disorder types, enhancing the sensitivity to pathological pronunciation patterns, reducing the risk of missed detections and misjudgments by general models, and outputting structured articulation marking results that intuitively present abnormal characters and their corresponding states, efficiently supporting clinical assessment and personalized rehabilitation plan development, and significantly improving the accuracy, targeting, and clinical translation efficiency of articulation disorder localization.

[0077] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0078] This application also provides an electronic device, including: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the various processes of the above-described articulation anomaly localization method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0079] For example, Figure 3 A schematic diagram of the physical structure of an electronic device is shown. (For example...) Figure 3As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communications bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other through the communications bus 340. The processor 310 can call logical instructions in the memory 330. The processor 310 is used to perform the following steps: acquiring audio and video data recorded when the target object speaks; performing text conversion and correction processing on the audio data in the audio and video data to obtain text data; performing lip region recognition and extraction on the video data in the audio and video data to obtain lip video data; extracting the audio segment and video segment corresponding to each character in the text data from the audio data and the lip video data to obtain data to be located; inputting the data to be located into the audio encoder and video encoder included in the multimodal large language model fine-tuned by the target training data, and performing cross-modal fusion through a unified embedding space to obtain articulation marking results, wherein the articulation marking results include the articulation state corresponding to each character in the text data; the target training data includes multiple sample audio and video data labeled with character-level articulation states, wherein the recording object of the sample audio and video data and the target object belong to the same articulation disorder type. The processor 310 can also execute other schemes in the embodiments of this application, which will not be further described here.

[0080] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0081] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described speech anomaly localization method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0082] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0083] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0084] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0085] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0086] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0087] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0088] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0089] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0090] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0091] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for locating articulation abnormalities, characterized in that, include: Acquire audio and video data recorded while the target person is speaking; The audio data in the audio and video data is subjected to text conversion and correction processing to obtain text data; The lip region is identified and extracted from the video data in the audio and video data to obtain lip video data; The audio segment and video segment corresponding to each character in the text data are extracted from the audio data and the lip video data to obtain the data to be located. The data to be located is input into the audio encoder and video encoder of the multimodal large language model fine-tuned by the target training data. Cross-modal fusion is performed through a unified embedding space to obtain the articulation marking result. The articulation marking result includes the articulation state corresponding to each character in the text data. The target training data includes multiple sample audio and video data labeled with character-level articulation states. The recording objects of the sample audio and video data and the target object belong to the same articulation disorder type.

2. The method for locating articulation abnormalities according to claim 1, characterized in that, Extracting the audio segment and video segment corresponding to each character in the text data from the audio data and the lip video data includes: The audio data and the lip video data are time-aligned; The audio data and the text data are aligned at the character level to determine the audio segment corresponding to each character in the text data, as well as the start and end times of the audio segment. For each character in the text data, the corresponding video segment is extracted from the lip video data based on the start and end times of the character.

3. The method for locating articulation abnormalities according to claim 1, characterized in that, The data to be located is input into the audio encoder and video encoder of a multimodal large language model fine-tuned with target training data. Cross-modal fusion is performed through a unified embedding space to obtain articulation marker results, including: Feature extraction is performed on the audio and video segments corresponding to each character in the data to be located, to obtain the audio feature vector and visual feature vector corresponding to each character; The audio feature vector and visual feature vector corresponding to each character are concatenated with the modal marker, and the concatenation result is input into the multimodal large language model. Obtain the pronunciation type embedding vector corresponding to each character output by the multimodal large language model; Based on the pronunciation type embedding vector corresponding to each character, and using the articulation state classification module, the articulation state corresponding to each character is determined, and the articulation marking result is obtained.

4. The method for locating articulation abnormalities according to claim 1, characterized in that, The audio data in the audio and video data undergoes text conversion and correction processing to obtain text data, including: Obtain the script containing the content of the target object's speech; The audio data is preprocessed and then converted into text to obtain the original text data; The original text data is matched and corrected based on the content script to obtain the text data.

5. The method for locating articulation abnormalities according to claim 1, characterized in that, The video data in the audio and video data is subjected to lip region recognition and extraction to obtain lip video data, including: Perform facial recognition on the video data to obtain facial video data; The key points of the lip region are located and extracted from the face video data to obtain the lip video data.

6. The method for locating articulation abnormalities according to claim 1, characterized in that, After obtaining the articulation marker results, the method further includes: Based on the articulation marking results, obtain the erroneous words with abnormal articulation states, the articulation abnormality type corresponding to the erroneous words, and their occurrence positions in the audio and video data; Obtain correction suggestions that match the articulation anomaly type of the erroneous word; Based on the type of articulation error, location of occurrence, and correction suggestions corresponding to the erroneous word, an articulation error location report is generated.

7. The method for locating articulation abnormalities according to claim 6, characterized in that, After obtaining the articulation marker results, the method further includes: On the articulation error location display interface, the error words with an abnormal articulation status are displayed; Upon receiving a playback request for the erroneous word, the video and audio segments corresponding to the erroneous word are played.

8. A device for locating articulation abnormalities, characterized in that, include: The first acquisition module is used to acquire audio and video data recorded when the target object speaks; The first processing module is used to perform text conversion and correction processing on the audio data in the audio and video data to obtain text data; The second processing module is used to perform lip region recognition and extraction on the video data in the audio and video data to obtain lip video data. The extraction module is used to extract the audio segment and video segment corresponding to each character in the text data from the audio data and the lip video data to obtain the data to be located; The articulation localization module is used to input the data to be localized into the audio encoder and video encoder of the multimodal large language model fine-tuned by the target training data, and perform cross-modal fusion through a unified embedding space to obtain articulation marking results. The articulation marking results include the articulation state corresponding to each character in the text data. The target training data includes multiple sample audio and video data labeled with character-level articulation states. The recording objects of the sample audio and video data and the target object belong to the same articulation disorder type.

9. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the articulation anomaly localization method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the articulation anomaly localization method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Visual face contour motion-based dysarthria speech recognition method and system

    CN113241065A

  • Speech and language disorder assessment and rehabilitation training system based on artificial intelligence

    CN117038071A

  • Personalized customized dysarthria speech recognition method and system

    CN118865958A

  • Seat system for airplane

    KR102717699B1

  • Automated recommendation tool to improve intelligiblity in speech dysarthria

    US20250299595A1