Speech recognition method and device for cognitive competence assessment of old people

By performing Mel-spectrum processing and cross-modal fusion analysis on the speech signals of the elderly, the problem of high misjudgment rate caused by single-dimensional analysis in existing technologies has been solved, and accurate assessment of the cognitive abilities of the elderly has been achieved.

CN121483286APending Publication Date: 2026-02-06BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511815549.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing technologies rely on single-dimensional analysis in assessing the cognitive abilities of the elderly, lacking cross-modal information fusion and verification, resulting in a high misjudgment rate and failing to meet the needs of accurate assessment.

Method used

By collecting speech signals from elderly people, generating audio signals and performing Mel spectrum processing, identifying speech segments and silence segments, extracting audio feature data, decoding the audio signals into text data, and combining the text data with a preset model to identify emotions and cognitive abilities, a comprehensive analysis result is generated, achieving cross-modal fusion of acoustic features, audio features and text semantics.

Benefits of technology

It reduces the misjudgment rate of cognitive state assessment, outputs more accurate cognitive state assessment results, and provides a multi-dimensional data fusion mechanism to avoid the one-sidedness of a single data dimension.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121483286A_ABST
    Figure CN121483286A_ABST
Patent Text Reader

Abstract

The invention discloses a speech recognition method and device for cognitive competence assessment of old people, and relates to the technical field of speech recognition, and the method comprises the steps: collecting the speech of the old people during speaking, and generating an audio signal; processing the audio signal to obtain a Mel spectrum, and recognizing a voice segment and a mute segment based on the audio signal and a preset voice threshold value as a voice recognition result; performing feature extraction on the audio signal to obtain audio feature data; decoding the audio signal into text data; obtaining a semantic recognition result based on the text data, a preset text recognition model, the voice recognition result and the audio feature data; performing integration processing based on the voice recognition result, the audio feature data, the semantic recognition result and the Mel spectrum to generate a comprehensive analysis result for evaluating the cognitive competence of the old people; according to the multi-dimensional data fusion mechanism, the acoustic features, the audio features and the text semantics can be fused in a cross-modal mode, and the one-sidedness of a single data dimension is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech recognition technology, and in particular to a speech recognition method and apparatus for assessing the cognitive abilities of the elderly. Background Technology

[0002] With the increasing aging of the global population, the need for early screening and assessment of cognitive impairment in the elderly is becoming increasingly urgent. Traditional assessment methods mainly rely on manual questionnaires, which are inefficient, time-consuming, and highly subjective. In addition, some products attempt to provide single-dimensional auxiliary assessments through audio or text. For example, they use general speech recognition models to convert the elderly's speech into text and then use simple keyword matching to determine their cognitive status; or they extract basic parameters such as pitch and speech rate from audio devices as reference indicators of cognitive function.

[0003] Currently, although existing technologies have initially laid the foundation for intelligence, they still have obvious limitations. Because they rely on only a single dimension of text or audio for analysis and lack cross-modal information fusion and verification, existing technologies have a high misjudgment rate in cognitive state judgment and are difficult to meet the actual use needs of accurate assessment. Summary of the Invention

[0004] The purpose of this application is to provide a speech recognition method and device for assessing the cognitive abilities of the elderly, which can solve the problem of relying on only a single dimension for analysis and lacking cross-modal information fusion and verification.

[0005] To achieve the above objectives, this application provides the following solution: In a first aspect, this application provides a speech recognition method for assessing the cognitive abilities of older adults, comprising: Collect the speech of elderly people and generate audio signals; The audio signal is processed to obtain the Mel spectrum, and speech segments and silence segments are identified based on the audio signal and a preset speech threshold as the speech recognition result; Feature extraction is performed on the audio signal to obtain audio feature data; Decode the audio signal into text data; The semantic recognition result is obtained based on the text data, the preset text recognition model, the speech recognition result, and the audio feature data; The speech recognition results, audio feature data, semantic recognition results, and Mel spectrum are integrated and processed to generate a comprehensive analysis result for assessing the cognitive abilities of the elderly person.

[0006] In one embodiment, the step of identifying speech segments and silence segments based on the audio signal and a preset speech threshold specifically includes: In the audio signal, segments with energy below the speech threshold are identified as silent segments, and segments with energy above or equal to the speech threshold are identified as speech segments.

[0007] In one embodiment, the audio feature data includes pitch data, pitch vibrato data, and energy data; The pitch data includes at least the fundamental frequency mean, peak pitch, minimum pitch, and pitch range; the pitch jitter data includes at least the fundamental frequency standard deviation, fundamental frequency jitter, and amplitude jitter; and the energy data includes at least the energy mean, peak energy, and energy range.

[0008] In one embodiment, the step of obtaining the semantic recognition result based on the text data, the preset text recognition model, the speech recognition result, and the audio feature data specifically includes: Based on the text data and the preset text recognition model, emotion recognition and cognitive ability recognition are performed to obtain emotion recognition results and cognitive ability recognition results. The semantic recognition result is obtained based on the emotion recognition result, the cognitive ability recognition result, the speech recognition result, and the audio feature data.

[0009] In one embodiment, the step of performing emotion recognition specifically includes: The number of various emotional terms in the text data is counted, and the percentage of each emotional term to the total number of emotions is generated.

[0010] In one embodiment, the step of recognizing cognitive abilities specifically includes: The statistics include the number of pauses, the number of words, the number of logical problems, the number of word errors, the number of repeated words, and the speaking speed of the text data.

[0011] Secondly, this application also provides a speech recognition device for assessing the cognitive abilities of the elderly, comprising: The voice data acquisition module collects the voice of elderly people when they speak and generates audio signals; The acoustic feature extraction and analysis module processes the audio signal to obtain the Mel spectrum, and identifies speech segments and silence segments based on the audio signal and a preset speech threshold as the speech recognition result; The audio feature analysis module extracts features from the audio signal to obtain audio feature data; The speech-to-text module decodes the audio signal into text data; The semantic analysis module obtains semantic recognition results based on the text data, the preset text recognition model, the speech recognition results, and the audio feature data; The comprehensive analysis module integrates and processes the speech recognition results, audio feature data, semantic recognition results, and Mel spectrum to generate a comprehensive analysis result for assessing the cognitive abilities of the elderly person.

[0012] Thirdly, this application also provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method.

[0013] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program implements the above-described method when executed by a processor.

[0014] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0015] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a speech recognition method for assessing the cognitive abilities of the elderly. It processes audio signals to obtain a Mel spectrum, and identifies speech segments and silence segments based on the audio signal and a preset speech threshold as the speech recognition result. Features are extracted from the audio signal to obtain audio feature data, which is then decoded into text data. Semantic recognition results are obtained based on the text data, a preset text recognition model, the speech recognition result, and the audio feature data. Finally, the speech recognition result, audio feature data, semantic recognition result, and Mel spectrum are integrated to generate a comprehensive analysis result for assessing the cognitive abilities of the elderly. Thus, this application's multi-dimensional data fusion mechanism can cross-modally fuse acoustic features, audio features, and text semantics, avoiding the limitations of a single data dimension, effectively reducing the misjudgment rate in cognitive state assessment, and outputting more accurate cognitive state assessment results. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart of a speech recognition method for assessing the cognitive abilities of the elderly, according to an embodiment of this application. Figure 2 This is a flowchart of a speech recognition system for assessing the cognitive abilities of the elderly, according to an embodiment of this application. Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0020] See Figure 1 This application provides a speech recognition method for assessing the cognitive abilities of the elderly, comprising the following steps: S100: Collects the speech of elderly people when they speak and generates audio signals; S200: Process the audio signal to obtain the Mel spectrum, and identify speech segments and silence segments based on the audio signal and a preset speech threshold, as the speech recognition result; S300: Extracts features from audio signals to obtain audio feature data; S400: Decodes audio signals into text data; S500: Semantic recognition results are obtained based on text data, a preset text recognition model, speech recognition results, and audio feature data; S600: Based on speech recognition results, audio feature data, semantic recognition results, and Mel spectrum, it integrates and processes these data to generate comprehensive analysis results for assessing the cognitive abilities of older adults.

[0021] In the S100, the voice acquisition terminal responds to the elderly person's trigger commands, such as pressing the record button to activate the built-in microphone and record voice data of a preset duration according to preset parameters, for example, recording 3 minutes of voice data as a WAV audio file. After recording, the voice acquisition terminal uploads the audio file to the server via a wireless network, and the server receives and stores the file. This step obtains raw voice data of standardized duration, providing a complete sample basis for subsequent analysis.

[0022] For example, the voice recording terminal can be a portable recording device such as a smartphone, tablet, or smartwatch. The device has a built-in high-sensitivity noise-canceling microphone that can respond to physical button triggers, touchscreen commands, or voice wake-up commands from the elderly, and complete voice recording and uploading functions. The high-sensitivity noise-canceling microphone is adapted to the low-volume speaking characteristics of the elderly, and the 3-minute timed recording is suitable for the sustained attention span of the elderly.

[0023] In S200, the steps for processing the audio signal to obtain the Mel spectrum are as follows: pre-emphasis is applied to the audio signal; the pre-emphasis processed audio signal is divided into frames with a frame length of 0.025 seconds and a frame shift of 0.01 seconds, and a Hanning window is applied to each frame signal to reduce spectral leakage; a 512-point Fourier transform is performed on each windowed frame signal to calculate the power spectrum; the power spectrum is processed through 40 Mel filter banks to obtain the Mel spectrum.

[0024] It should be noted that in the audio preprocessing stage, the pre-emphasis coefficient can be set to 0.95 and the frame length to 0.025s. This is specifically adjusted for the characteristics of elderly people's unclear pronunciation and weak breath, in order to improve the feature extraction effect of unclear speech.

[0025] The steps for identifying speech segments and silence segments based on audio signals and preset speech thresholds specifically include: in the audio signal, segments with energy below the speech threshold are identified as silence segments, and segments with energy above or equal to the speech threshold are identified as speech segments. The server stores Mel-spectrum data and speech recognition results in a local database.

[0026] For example, a duration of 3-5 seconds can be defined as a cognitive lapse in older adults. Specifically, among identified silent segments, those long silences lasting between 3 and 5 seconds are identified as lapses associated with cognitive decline. These unusually long silences typically reflect difficulties older adults experience in organizing thoughts, searching for words, or performing expressions, and are important acoustic indicators for assessing their cognitive function.

[0027] For example, the speech threshold ranges from 0.001 to 0.003. Those skilled in the art can set specific values ​​for the speech threshold according to actual conditions. Specifically, during speech processing, the speech threshold value can be dynamically adjusted based on the level of ambient background noise. When noise is high, the threshold is adaptively increased to effectively filter background noise; when noise is low, the threshold is correspondingly decreased to accurately capture weak speech, thereby achieving speech segmentation under different noise environments.

[0028] In the S300, feature extraction is performed on the audio signal to obtain audio feature data. The specific process is as follows: The pydub library is used to unify MP3 format to WAV format, as WAV format is more suitable for the opensmile toolkit. Frame-level feature extraction is employed, using opensmile toolkit settings to extract basic acoustic features (pitch jitter data and energy-related data) for each frame (default 10ms / frame): such as jitter, shimmer, F0 (fundamental frequency), and RMS energy. Functional-level feature extraction is used, performing global statistics on frame-level features, such as statistical mean, standard deviation, maximum / minimum values, and range (difference between maximum and minimum values). Additional statistical calculations are performed on frame-level features to compensate for deficiencies in functional-level features, such as the distribution characteristics of F0.

[0029] Specifically, the audio feature data includes pitch data, pitch tremolo data, and energy data; the pitch data includes at least the fundamental frequency mean, peak pitch, minimum pitch, and pitch range; the pitch tremolo data includes at least the fundamental frequency standard deviation (reflecting the degree of pitch fluctuation), fundamental frequency jitter, and amplitude jitter; and the energy data includes at least the energy mean, peak energy, and energy range.

[0030] In the S400, Automatic Speech Recognition (ASR) technology is used to decode audio signals and generate text data.

[0031] In this step, automatic speech recognition technology uses components such as acoustic models and language models to convert the speech content in the input audio signal into a text sequence that can be recognized and processed by a computer, providing a data foundation for subsequent semantic analysis.

[0032] In S500, the steps of obtaining semantic recognition results based on text data, a preset text recognition model, speech recognition results, and audio feature data specifically include: performing emotion recognition and cognitive ability recognition based on text data and a preset text recognition model to obtain emotion recognition results and cognitive ability recognition results; and obtaining semantic recognition results based on emotion recognition results, cognitive ability recognition results, speech recognition results, and audio feature data.

[0033] The preset text recognition model is a specially fine-tuned Large Language Model (LLM). It is a dedicated model that is pre-trained on the basis of a general Large Language Model using annotated corpus specifically related to cognitive and emotion analysis of the elderly. It can more accurately identify features related to cognitive state in the language of the elderly, so that the analysis results are more in line with the language characteristics and assessment needs of the elderly population.

[0034] In this step, the emotion recognition process specifically includes: counting the number of various emotion words in the text data and generating the percentage of each emotion word to the total number of emotions.

[0035] A LoRa fine-tuning technique is employed, inputting a large amount of labeled data into a large model for fine-tuning training. This labeled data clearly defines the emotional category (e.g., positive, negative, optimistic, or pessimistic) of words or sentences, enabling the large model to better capture the semantic relationships between words and the semantic information of sentences. Through vector mapping, sentences or words are mapped to their corresponding emotional states. Furthermore, the number of positive and negative words is obtained separately through detection. The positive percentage is calculated as: the number of positive words divided by the sum of the number of positive and negative words.

[0036] For example, count the number of positive words (e.g., "happy" and "satisfied") and negative words (e.g., "sad" and "angry") in the text, and calculate the proportion of positive words and negative words to the total number of words, respectively.

[0037] The number of optimistic words (e.g., “hope” and “it will be alright”) and pessimistic words (e.g., “despair” and “useless”) in the text is counted, and the proportion of optimistic words and pessimistic words to the total number of words is calculated separately.

[0038] Identify seven emotions contained in the text, such as anger, disgust, sadness, joy, neutrality, surprise, and fear. Count the number of times each emotion appears and calculate its percentage in the total number of identified emotions (number of times each emotion / total number of emotions).

[0039] This step, which involves recognizing cognitive abilities, specifically includes: counting the number of pauses, the number of words, the number of logical problems, the number of word errors, the number of repeated words, and the speech rate in the text data.

[0040] For example, the number of pauses / interruptions in text is counted. This is primarily based on two types of features: repeated interjections (such as "um" or "ah") and incomplete breaks in sentences. The pre-defined text recognition model effectively identifies these features, marking each interjection as a pause and treating structurally incomplete sentences as signals of potential cognitive shifts or expression blockages, thus more accurately counting the frequency of pauses.

[0041] Count the number of characters / words in the text. The number of characters is calculated by counting the number of tokens, generally one token corresponds to one character.

[0042] The system counts the number of logical problems, such as contradictory statements and incorrect causal relationships. It uses deep learning models to perform contextual analysis and intent recognition on the text, and leverages big data and historical knowledge for multi-dimensional reasoning and analysis to determine if logical contradictions or illogical reasoning exist within the text.

[0043] The system counts the number of incorrect word choices, such as inappropriate word usage and typos. By learning from a large amount of correct text data, it masters the correct usage and context of vocabulary. When processing input text, it analyzes the semantics, grammar, and context of words to determine whether a word is used appropriately.

[0044] This involves counting the number of times a word is repeated, for example, the number of times the same word appears meaninglessly in a text. It primarily uses natural language processing and deep learning techniques to determine the meaning of words in the text. The text is parsed into a sequence of tokens that the model can process, and then these tokens are analyzed using a deep learning model, combined with contextual information and semantic understanding techniques, to determine the meaning of each word in the text.

[0045] A speech rate test is conducted. Speech is converted to text using speech recognition technology, and the duration of the speech is recorded. The total number of words in the text is counted, and the average speech rate (unit: words / second) is calculated by dividing the total number of words by the speech duration. The actual average speech rate is compared with a standard range (80-160 words / minute) to determine if the speech rate is normal. For example, an actual speech rate below 80 words / minute is considered slow, 80-160 words / minute is normal, and an actual speech rate above 160 words / minute is considered fast.

[0046] For audio feature time points and corresponding speech-to-text time points, pitch data is more likely to correspond to confidence and a positive state, while lower pitch is the opposite. Similarly, pitch tremor data can help better analyze emotional or cognitive states; for example, significant audio tremor indicates that the person may be in an excited state. Combining this with text can better predict trends and improve accuracy. Changes in energy-related data better reflect the level of effort a person exerts when speaking; higher energy may indicate a confident speech or high spirits, while lower energy may indicate fatigue, illness, or low mood.

[0047] A preliminary interpretation can be made by combining the emotion recognition results, cognitive ability recognition results, speech recognition results, and audio feature data obtained from the above steps: For example, a noticeable tremor in pitch and the presence of multiple sad words in the text may reflect a depressed mood.

[0048] For example, if negative words appear in the text, it may initially be judged to be a sad emotion. Combining the low mean fundamental frequency and low energy peak of the audio feature data, which reflect the low tone of the voice, and the identified silent segments, which correspond to pauses when the emotion is low, the sadness emotion can be further confirmed.

[0049] For example, if the text contains interjections such as "um" or "ah", it is initially identified as a pause or stutter. If the corresponding time period is a silent segment and the pitch jitter value is abnormally high, the pause can be considered as a cognitive-related stutter or an abnormal tone pause, thus improving the recognition accuracy of cognitive ability.

[0050] In S600, for example, according to preset clinical cognitive assessment criteria, corresponding weights are assigned to each data item in speech recognition results, audio feature data, semantic recognition results, and acoustic features reflected by Mel spectrum; based on the assigned weights, the data items are weighted and integrated to obtain a comprehensive analysis result.

[0051] For example, clinical cognitive assessment criteria refer to the experience and rules regarding the importance of various cognitive indicators extracted from mature, clinically recognized cognitive assessment scales and medical knowledge. That is, in this application, which features are more indicative of cognitive impairment and thus carry a higher weight in the overall score?

[0052] Clinical cognitive assessment standards mainly include standardized scale assessments, specific cognitive domain tests, and clinical grading standards. Commonly used tools include the Montreal Cognitive Assessment Scale (MoCA) and the Mini-Mental State Examination (MMSE).

[0053] The comprehensive analysis results represent a multi-dimensional quantitative assessment of the cognitive abilities of older adults. By weighted fusion of multimodal data, including acoustic features, audio characteristics, and text semantics, a structured comprehensive report is generated, intuitively reflecting the cognitive state of the participants and providing objective evidence for professional judgment. The server stores the comprehensive report in the user database and can push it to terminal devices for viewing via an interface.

[0054] When generating a comprehensive report, different assessment indicators have varying degrees of importance in determining the final cognitive ability. Weighting quantifies this importance, and its specific value is set based on clinical cognitive assessment standards and research findings. For example, clinical cognitive assessment standards indicate that certain indicators (such as abnormal speech rate or logical inconsistency) are more strongly associated with cognitive impairment and therefore receive higher weightings in the comprehensive score, having a greater impact on the final result; other indicators may serve as supplementary references and thus have lower weightings.

[0055] In a specific example, an audio signal of an elderly person speaking was collected. The speech recognition results showed: a 4-second silence segment was detected, and the effective speech segment was fragmented by multiple short silences. Mel spectrum analysis revealed that during the 4-second silence segment, the spectral energy remained close to zero, and the dynamic range of the speaking portion was narrow, indicating insufficient vocal vitality. Audio feature data showed a slow overall speech rate, low average energy, and noticeable pitch tremolo before pauses. Semantic recognition results indicated that the converted text contained multiple repetitive words and incomplete sentences.

[0056] Furthermore, different weights were assigned to the abnormalities of the above-mentioned different characteristics to obtain a comprehensive analysis conclusion. The resulting report pointed out that when the subjects recalled and described simple events, they exhibited significant thought interruptions (long silences), speech disfluency (repetitions, incomplete sentences), and decreased vocal energy (low energy, narrow spectrum). These multimodal evidences jointly suggest that they have the risk of difficulty in working memory retrieval and decline in executive function, and further clinical cognitive assessment is recommended.

[0057] Based on the same inventive concept, this application also provides a speech recognition device for assessing the cognitive abilities of the elderly. The solution provided by this device is similar to the solution described in the above-described method. Therefore, the specific limitations of one or more embodiments of the speech recognition device for assessing the cognitive abilities of the elderly provided below can be found in the limitations of the speech recognition method for assessing the cognitive abilities of the elderly described above, and will not be repeated here.

[0058] See Figure 2 This application also includes a speech recognition device for assessing the cognitive abilities of older adults, comprising: The voice data acquisition module collects the voice of elderly people when they speak and generates audio signals; The acoustic feature extraction and analysis module processes the audio signal to obtain the Mel spectrum, and identifies speech segments and silence segments based on the audio signal and preset speech thresholds as the speech recognition result; The audio feature analysis module extracts features from the audio signal to obtain audio feature data; The speech-to-text module decodes audio signals into text data; The semantic analysis module obtains semantic recognition results based on text data, a preset text recognition model, speech recognition results, and audio feature data. The comprehensive analysis module integrates and processes speech recognition results, audio feature data, semantic recognition results, and Mel spectrum to generate comprehensive analysis results for assessing the cognitive abilities of older adults.

[0059] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 3As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and databases. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media to run. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection.

[0060] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0061] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0062] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0063] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0064] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0065] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0066] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0067] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0068] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A speech recognition method for assessing the cognitive abilities of the elderly, characterized in that, include: Collect the speech of elderly people and generate audio signals; The audio signal is processed to obtain the Mel spectrum, and speech segments and silence segments are identified based on the audio signal and a preset speech threshold as the speech recognition result; Feature extraction is performed on the audio signal to obtain audio feature data; Decode the audio signal into text data; The semantic recognition result is obtained based on the text data, the preset text recognition model, the speech recognition result, and the audio feature data; The speech recognition results, audio feature data, semantic recognition results, and Mel spectrum are integrated and processed to generate a comprehensive analysis result for assessing the cognitive abilities of the elderly person.

2. The speech recognition method for assessing the cognitive abilities of the elderly according to claim 1, characterized in that, The step of identifying speech segments and silence segments based on the audio signal and a preset speech threshold specifically includes: In the audio signal, segments with energy below the speech threshold are identified as silent segments, and segments with energy above or equal to the speech threshold are identified as speech segments.

3. The speech recognition method for assessing the cognitive abilities of the elderly according to claim 1, characterized in that, The audio feature data includes pitch data, pitch vibrato data, and energy data; The pitch data includes at least the fundamental frequency mean, peak pitch, minimum pitch, and pitch range; the pitch jitter data includes at least the fundamental frequency standard deviation, fundamental frequency jitter, and amplitude jitter; and the energy data includes at least the energy mean, peak energy, and energy range.

4. The speech recognition method for assessing the cognitive abilities of the elderly according to claim 1, characterized in that, The step of obtaining the semantic recognition result based on the text data, the preset text recognition model, the speech recognition result, and the audio feature data specifically includes: Based on the text data and the preset text recognition model, emotion recognition and cognitive ability recognition are performed to obtain emotion recognition results and cognitive ability recognition results. The semantic recognition result is obtained based on the emotion recognition result, the cognitive ability recognition result, the speech recognition result, and the audio feature data.

5. The speech recognition method for assessing the cognitive abilities of the elderly according to claim 4, characterized in that, The steps involved in emotion recognition include: The number of various emotional terms in the text data is counted, and the percentage of each emotional term to the total number of emotions is generated.

6. The speech recognition method for assessing the cognitive abilities of the elderly according to claim 5, characterized in that, The steps for cognitive ability identification specifically include: The statistics include the number of pauses, the number of words, the number of logical problems, the number of word errors, the number of repeated words, and the speaking speed of the text data.

7. A speech recognition device for assessing the cognitive abilities of the elderly, characterized in that, include: The voice data acquisition module collects the voice of elderly people when they speak and generates audio signals; The acoustic feature extraction and analysis module processes the audio signal to obtain the Mel spectrum, and identifies speech segments and silence segments based on the audio signal and a preset speech threshold as the speech recognition result; The audio feature analysis module extracts features from the audio signal to obtain audio feature data; The speech-to-text module decodes the audio signal into text data; The semantic analysis module obtains semantic recognition results based on the text data, the preset text recognition model, the speech recognition results, and the audio feature data; The comprehensive analysis module integrates and processes the speech recognition results, audio feature data, semantic recognition results, and Mel spectrum to generate a comprehensive analysis result for assessing the cognitive abilities of the elderly person.

8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the speech recognition method for assessing cognitive abilities of the elderly as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the speech recognition method for assessing the cognitive abilities of the elderly as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the speech recognition method for assessing the cognitive abilities of the elderly as described in any one of claims 1-6.