An interactive robot for hemiplegia and speech disorders

Through the Fourier image processing and repeated word recognition network, the problem of unclear speech in patients with hemiplegia in limbs is solved, accurate speech recognition and sentence completion are achieved, and the clarity of communication is improved.

CN119181360BActive Publication Date: 2025-05-16THE SECOND AFFILIATED HOSPITAL OF GUANGZHOU MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411243320.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2025-05-16
Estimated Expiration
2044-09-05

AI Technical Summary

Technical Problem

Patients with hemiplegia experience unclear speech and difficulty in semantic recognition during communication, and cannot accurately identify complete sentences and continuous repetitive words.

Method used

Through speech Fourier image processing, the pronunciation set and time points in the speech signal are detected, missing words are predicted, and repeated word recognition networks are used to identify repeated or different words, predictive statements are constructed, and outputted through the sound playback device.

Benefits of technology

It improves the accuracy of speech recognition for patients with hemiplegia in limbs, predicts and completes missing words in sentences and skips repeated words to achieve more accurate semantic recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119181360B_ABST
    Figure CN119181360B_ABST
Patent Text Reader

Abstract

The invention discloses an interactive robot for hemiplegia and speech disorders. It performs pronunciation, finds corresponding words, and establishes an adjacent pronunciation relationship to find words with similar pronunciations, thereby preventing the situation where the user wants to say a word but pronounces it vaguely. These words are matched, and the words with blank pauses are predicted to prevent the user from having weak pronunciation or swallowing pronunciation due to physical reasons, thereby predicting the missing part of a paragraph. Repeated words are skipped according to a repeated word recognition network. The repeated word recognition network adopts the parameters in the modified threshold structure to train a repeated word recognition network, and the parameters on different networks are retained to achieve the purpose of setting networks with different structures. And the characteristics and predictions of the input neural network of the skipped repeated words are used to determine a better matching method for the sentence. Thereby achieving a more accurate technical effect of predicting the voice of a user with hemiplegia and speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an interactive robot for patients with limb hemiplegia and speech disorders. Background Art

[0002] At present, when users with hemiplegia and speech are communicating, they usually cannot speak a complete sentence, or the pronunciation of words between sentences is vague and not accurate. Or because they cannot speak a complete paragraph, there are pauses between sentences, and words are repeated continuously. This makes it difficult to recognize semantics. It is difficult to accurately recognize the voice of users with hemiplegia and speech. Summary of the invention

[0003] The purpose of the present invention is to provide an interactive robot for hemiplegia and speech disorders to solve the above problems existing in the prior art.

[0004] An embodiment of the present invention provides an interactive robot for hemiplegia and speech disorders, comprising a processor;

[0005] The processor is used to perform the following steps:

[0006] A speech signal and a speech Fourier image are obtained; the speech signal is speech with a language barrier received by a sound receiving device; the speech Fourier image represents a Fourier image converted from the speech signal;

[0007] Based on the speech Fourier image, corresponding pronunciations in the speech signal are detected, similar pronunciations are found, and multiple pronunciation sets and corresponding pronunciation time points are obtained; one pronunciation time point corresponds to one pronunciation set; and the pronunciation set contains similar pronunciations;

[0008] Based on the speech Fourier image, multiple pronunciation sets and corresponding pronunciation time points, predict missing words to obtain a prediction word matrix and a detection word matrix;

[0009] Using a repeated word recognition network, based on multiple pronunciation sets and a detection word matrix, repeated words or words with the same pronunciation but different meanings in a speech signal are detected to obtain an original matching word matrix; the original matching word matrix represents multiple words constituting a speech;

[0010] According to the predicted word matrix and the original matching word matrix, determining whether to add the predicted word, predicting the meaning of the sentence, and obtaining the predicted sentence;

[0011] The predicted sentence is converted into a play voice signal and sent to a sound playing device.

[0012] Optionally, the repeated word recognition network includes multiple layers of neurons;

[0013] Each layer of neurons is controlled by a threshold structure; n layers of neurons correspond to n threshold structures;

[0014] Each layer of neurons is fully connected to all other layers of neurons.

[0015] Optionally, the method of detecting repeated words or words with the same pronunciation but different meanings in the speech signal through a repeated word recognition network based on multiple pronunciation sets and a detection word matrix to obtain an original matching word matrix includes:

[0016] The detection word matrix is ​​divided into columns to obtain a plurality of word vectors; one word vector corresponds to one pronunciation time point; an element in the word vector represents a word corresponding to the pronunciation corresponding to one pronunciation time point;

[0017] According to the plurality of pronunciation sets, detecting whether pronunciation sets at adjacent pronunciation time points are the same pronunciation set;

[0018] If the pronunciation sets of adjacent pronunciation time points are the same pronunciation set, the pronunciation time points corresponding to the corresponding word vectors are filled into the repeated pronunciation time set;

[0019] Taking the word vectors corresponding to the pronunciation time points in the repeated pronunciation time set in the multiple word vectors as repeated word vectors to obtain a repeated word vector set; the elements in the repeated word vector set are arranged from early to late according to the pronunciation time points;

[0020] Taking the word vectors corresponding to the pronunciation time points in the repeated pronunciation time set except for the word vectors in the multiple word vectors as non-repeated word vectors to obtain a non-repeated word vector set; the elements in the non-repeated word vector set are arranged from early to late according to the pronunciation time points;

[0021] Based on the repeated pronunciation time set, the repeated word vector set and the non-repeated word vector set, an original matching word matrix is ​​obtained through a repeated word recognition network.

[0022] Optionally, the obtaining of an original matching word matrix based on the repeated pronunciation time set, the repeated word vector set and the non-repeated word vector set through a repeated word recognition network includes:

[0023] The value in the threshold structure corresponding to all non-repeated word vectors in the non-repeated word vector set is set to 1; the value in the threshold structure being 1 indicates that only the parameters of neurons in the layer corresponding to the adjacent word vectors are retained for full connection;

[0024] The number of pronunciation time points in the repeated pronunciation time set is used as the number of repetitions;

[0025] Arrange the natural numbers between 0 and the sum of the number of repetitions plus 1 in descending order to obtain a set of repetition values;

[0026] Using a repeated word vector in the repeated word vector set as a subscript repeated word vector;

[0027] Taking the value in the set of repeated values ​​that is equal to the subscript of the subscript repeated word vector as the subscript repeated value;

[0028] Put the subscript repetition value into the subscript repetition word vector corresponding threshold structure;

[0029] Arrange the natural numbers between 0 and the sum of the value in the threshold structure plus 1 in descending order to obtain a set of connection layer numbers; the elements in the set of connection layer numbers represent the difference between the number of layers of neurons fully connected with the subscript repeated word vector and the number of layers of neurons corresponding to the subscript repeated word vector;

[0030] The parameters of the full connection between the neuron corresponding to an element in the connection layer number set and the neuron corresponding to the subscript repeated word vector are retained;

[0031] Multiple elements in the connection layer number set correspond to obtaining multiple repeated word recognition networks with different parameters;

[0032] Based on the repeated word recognition networks with multiple different parameters and the multiple word vectors, an original matching word matrix is ​​obtained.

[0033] Optionally, the repeated word recognition network based on the multiple different parameters and the multiple word vectors obtains an original matching word matrix, including:

[0034] Inputting the plurality of word vectors into a repeated word recognition network to obtain a matching judgment value; a plurality of repeated word recognition networks with different parameters correspondingly obtain a plurality of matching judgment values;

[0035] Taking a value among multiple matching judgment values ​​that is greater than other matching judgment values ​​as a prediction matrix value;

[0036] The word vectors corresponding to the repeated word recognition network in the prediction matrix value are arranged in chronological order from early to late to obtain the original matching word matrix.

[0037] Optionally, predicting missing words based on the speech Fourier image, multiple pronunciation sets and corresponding pronunciation time points to obtain a predicted word matrix and a detected word matrix includes:

[0038] According to the pronunciation set, find the character corresponding to each pronunciation to obtain a character set; obtain multiple character sets corresponding to multiple pronunciation time points;

[0039] According to the pronunciation time points from early to late, the corresponding word sets are arranged and filled into a two-dimensional matrix in sequence to obtain a detection word matrix; the rows of the detection word matrix represent the corresponding pronunciation time points; the columns of the detection word matrix represent the word set corresponding to a pronunciation time point;

[0040] Based on the speech Fourier image and the detection word matrix, a blank word matrix and a blank pronunciation time point are obtained; the elements of the column corresponding to the blank pronunciation in the blank word matrix are 0;

[0041] Taking the pronunciation time points adjacent to the blank pronunciation time points in the blank word matrix as predicted blank pronunciation time points;

[0042] Perform word matching on the words corresponding to the predicted blank pronunciation time points in the blank word matrix to obtain a plurality of matching words;

[0043] According to the position of the characters in the matching word, a predicted blank character and a matching time point are obtained; the predicted blank character represents the characters in the matching word except the characters in the predicted blank pronunciation time point; the matching time point represents the blank pronunciation time point corresponding to the predicted blank character; one blank pronunciation time point corresponds to multiple predicted blank characters;

[0044] Fill the blank word matrix with the predicted blank words corresponding to the matching time point to obtain a predicted word matrix.

[0045] Optionally, obtaining a blank word matrix and blank pronunciation time points based on the speech Fourier image and the detection word matrix includes:

[0046] The position where the user's voice does not exist in the voice Fourier image is regarded as a blank pronunciation to obtain a blank pronunciation time point;

[0047] Obtain a blank vector; the blank vector is a column vector having the same number of columns as the detection word matrix and all element values ​​are 0;

[0048] According to the time sequence of the blank pronunciation time points and the pronunciation time points corresponding to the detection word matrix, a blank vector is inserted into the detection word matrix to obtain a blank word matrix.

[0049] Optionally, judging whether to add the predicted word according to the predicted word matrix and the original matching word matrix, predicting the meaning of the sentence, and obtaining the predicted sentence includes:

[0050] Matching the words in multiple columns in the predicted word matrix with the words in multiple columns in the original matching word matrix respectively to obtain multiple matching statements;

[0051] Inputting the matching sentence into a temporal convolutional network, predicting the fluency of the sentence, and obtaining a predicted sentence value;

[0052] The matching statement whose predicted statement value is greater than other predicted statement values ​​is taken as the predicted statement.

[0053] Optionally, based on the speech Fourier image, detecting corresponding pronunciations in the speech signal, finding similar pronunciations, and obtaining multiple pronunciation sets and corresponding pronunciation time points include:

[0054] Obtaining a plurality of constructed pronunciation sets; the constructed pronunciation sets comprising a plurality of similar pronunciations; the pronunciations representing the pronunciations corresponding to a character;

[0055] Obtaining a pronunciation time window; the pronunciation time window represents the length corresponding to a pronunciation;

[0056] Detecting the area corresponding to the pronunciation time window in the speech Fourier image to obtain a detected pronunciation;

[0057] Finding a corresponding constructed pronunciation set as the pronunciation set according to the detected pronunciation;

[0058] Multiple pronunciation time windows correspond to obtaining multiple pronunciation sets;

[0059] The starting time point of the time period corresponding to the pronunciation time window is taken as the pronunciation time point.

[0060] Optionally, the sound playing device is used to play voice.

[0061] Compared with the prior art, the embodiments of the present invention achieve the following beneficial effects:

[0062] The embodiment of the present invention also provides an interactive robot for hemiplegia and speech disorders, comprising a processor; the processor is used for the following steps: obtaining a speech signal and a speech Fourier image; the speech signal is a speech with a speech disorder received by a sound receiving device; the speech Fourier image represents the Fourier image converted from the speech signal; based on the speech Fourier image, detecting the corresponding pronunciation in the speech signal, finding similar pronunciations, and obtaining multiple pronunciation sets and corresponding pronunciation time points; one pronunciation time point corresponds to one pronunciation set; the pronunciation sets contain similar pronunciations; based on the speech Fourier image, predicting missing words, and obtaining a predicted word matrix and a detected word matrix; through a repeated word recognition network, based on multiple pronunciation sets and the detected word matrix, detecting repeated words or words with the same pronunciation but different meanings in the speech signal, and obtaining an original matching word matrix; the original matching word matrix represents multiple words constituting a speech; judging whether to add a predicted word based on the predicted word matrix and the original matching word matrix, predicting the meaning of a sentence, and obtaining a predicted sentence; converting the predicted sentence into a play speech signal, and sending it to a sound playing device.

[0063] The present invention performs pronunciation, finds corresponding words, and establishes an adjacent pronunciation relationship to find words with similar pronunciations, thereby preventing the user from saying a word but pronouncing it ambiguously. These words are matched, and the words with blank pauses are predicted to prevent the user from having weak pronunciation or swallowing pronunciation due to physical reasons, thereby predicting the missing part of a paragraph. Repeatedly spoken words are skipped according to a repeated word recognition network. The repeated word recognition network adopts the parameters in the modified threshold structure to train a repeated word recognition network, and the parameters on different networks are retained to achieve the purpose of setting networks with different structures. And the characteristics and predictions of the input neural network of the skipped repeated words are used to determine a better matching method for the sentence. Thereby achieving a more accurate technical effect of predicting the speech of a user with limb hemiplegia and language.

[0064] Figure 1 It is a method flow chart of an interactive robot for limb hemiplegia and speech disorders provided by an embodiment of the present invention.

[0065] Figure 2 It is a partial structural schematic diagram of a repeated word recognition network in an interactive robot for limb hemiplegia and speech disorders provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0066] The present invention will be described in detail below in conjunction with the accompanying drawings.

[0067] Example 1

[0068] like Figure 1 As shown, an embodiment of the present invention provides an interactive robot for hemiplegia and speech disorders, wherein the interactive robot for hemiplegia and speech disorders includes a processor;

[0069] The processor is used to perform the following steps:

[0070] S101: obtaining a speech signal and a speech Fourier image; the speech signal is speech with a speech barrier received by a sound receiving device; the speech Fourier image represents a Fourier image converted from the speech signal;

[0071] In this embodiment, the speech signal is converted into a Fourier image by Fourier transform.

[0072] S102: Based on the speech Fourier image, detect the corresponding pronunciation in the speech signal, find similar pronunciations, and obtain multiple pronunciation sets and corresponding pronunciation time points; one pronunciation time point corresponds to one pronunciation set; and the pronunciation set contains similar pronunciations;

[0073] S103: predicting missing words based on the speech Fourier image, multiple pronunciation sets and corresponding pronunciation time points, and obtaining a prediction word matrix and a detection word matrix;

[0074] S104: Detect repeated words or words with the same pronunciation but different meanings in the speech signal through a repeated word recognition network based on multiple pronunciation sets and a detection word matrix, and obtain an original matching word matrix; the original matching word matrix represents multiple words constituting a speech.

[0075] The detection word matrix is ​​sequentially input into each layer of the neural network in the repeated word recognition network by columns to extract features.

[0076] S105: judging whether to add the predicted word according to the predicted word matrix and the original matching word matrix, predicting the meaning of the sentence, and obtaining the predicted sentence.

[0077] Among them, semantic recognition is performed to obtain a predicted sentence.

[0078] S106: Convert the predicted sentence into a play voice signal and send it to a sound play device.

[0079] Optionally, the robot also includes: a sound receiving device and a sound playing device.

[0080] Optionally, the repeated word recognition network includes multiple layers of neurons;

[0081] Among them, the connection structure of the three layers in the repeated word recognition network is as follows Figure 2 shown.

[0082] Each layer of neurons is controlled by one threshold structure; n layers of neurons correspond to n threshold structures.

[0083] The threshold structure user controls whether each layer of neurons is multiplied by the parameters on its network;

[0084] Each layer of neurons is fully connected to all other layers of neurons.

[0085] Optionally, the method of detecting repeated words or words with the same pronunciation but different meanings in the speech signal through a repeated word recognition network based on multiple pronunciation sets and a detection word matrix to obtain an original matching word matrix includes:

[0086] The detection word matrix is ​​divided into columns to obtain a plurality of word vectors; one word vector corresponds to one pronunciation time point; and an element in the word vector represents a word corresponding to the pronunciation corresponding to one pronunciation time point.

[0087] According to the plurality of pronunciation sets, detecting whether pronunciation sets at adjacent pronunciation time points are the same pronunciation set;

[0088] The pronunciation sets at adjacent pronunciation time points may be two pronunciation sets to determine whether they are the same pronunciation set, or multiple pronunciation sets to determine whether they are the same pronunciation set.

[0089] If the pronunciation sets of adjacent pronunciation time points are the same pronunciation set, the pronunciation time points corresponding to the corresponding word vectors are filled into the repeated pronunciation time set;

[0090] The pronunciation time points in the repeated pronunciation time set are adjacent and the corresponding pronunciation sets are the same.

[0091] Wherein, the repeated pronunciation time set is one or more.

[0092] Taking the word vectors corresponding to the pronunciation time points in the repeated pronunciation time set in the multiple word vectors as repeated word vectors to obtain a repeated word vector set; the elements in the repeated word vector set are arranged from early to late according to the pronunciation time points;

[0093] Taking the word vectors corresponding to the pronunciation time points in the repeated pronunciation time set except for the word vectors in the multiple word vectors as non-repeated word vectors to obtain a non-repeated word vector set; the elements in the non-repeated word vector set are arranged from early to late according to the pronunciation time points;

[0094] Based on the repeated pronunciation time set, the repeated word vector set and the non-repeated word vector set, an original matching word matrix is ​​obtained through a repeated word recognition network.

[0095] Optionally, the obtaining of an original matching word matrix based on the repeated pronunciation time set, the repeated word vector set and the non-repeated word vector set through a repeated word recognition network includes:

[0096] The value in the threshold structure corresponding to all non-repeated word vectors in the non-repeated word vector set is set to 1; the value in the threshold structure being 1 indicates that only the parameters of neurons in the layer corresponding to the adjacent word vectors are retained for full connection;

[0097] Among them, the parameters when fully connected with other layers are ignored.

[0098] The number of pronunciation time points in the repeated pronunciation time set is used as the number of repetitions;

[0099] Arrange the natural numbers between 0 and the sum of the number of repetitions plus 1 in descending order to obtain a set of repetition value sets.

[0100] If the number of repetitions is 2, then 2+1=3, and the natural numbers between 0 and 3, namely 1 and 2, are taken as elements in the repetition value set. The repetition value set is {2, 1}.

[0101] The value in the repeated value set indicates the number of layers of neurons that a neuron corresponding to a repeated word vector will be connected to later. For example, if the first repeated word vector corresponds to the first layer of neurons and the corresponding repeated value is 2, then the first layer of neurons must retain the parameters for full connection with the second layer of neurons, and retain the parameters for full connection with the third layer of neurons.

[0102] Using a repeated word vector in the repeated word vector set as a subscript repeated word vector;

[0103] Taking the value in the set of repeated values ​​that is equal to the subscript of the subscript repeated word vector as the subscript repeated value;

[0104] Among them, if the subscript repeated word vector is the second element in the repeated word vector set, then the subscript is 1, and the value with subscript 1 in the repeated value set {2,1} is used as the subscript repeated value, that is, the subscript repeated value is 1.

[0105] The subscript repetition value is placed into the subscript repetition word vector corresponding threshold structure.

[0106] The threshold structure is used to save repeated subscript values ​​so as to determine whether to retain the parameters between neurons.

[0107] Wherein, one subscript repeated word vector corresponds to one threshold structure. The number of the threshold structures is the number of word vectors;

[0108] The natural numbers between 0 and the sum of the value in the threshold structure plus 1 are arranged in order from large to small to obtain a set of connection layer numbers; the elements in the set of connection layer numbers represent the difference between the number of layers of neurons that are fully connected with the subscript repeated word vector and the number of layers of neurons corresponding to the subscript repeated word vector.

[0109] The value of the element in the connection layer number set is the difference between the number of layers greater than the number of subscript repeated layers and the number of subscript repeated layers. The subscript repeated layers represent the number of layers corresponding to the neurons corresponding to the subscript repeated word vector.

[0110] The elements in the connection layer number set specifically indicate which layer of neurons are to be fully connected. For example, if the value of an element in the connection layer number set is 2, and the first repeated word vector corresponds to the first layer of neurons, it means that the parameters for fully connecting the first layer of neurons with the neurons of the corresponding layer separated by 2, that is, the third layer of neurons, are to be retained.

[0111] Based on the repeated word recognition networks with multiple different parameters and the multiple word vectors, an original matching word matrix is ​​obtained.

[0112] Optionally, the repeated word recognition network based on the multiple different parameters and the multiple word vectors obtains an original matching word matrix, including:

[0113] The multiple word vectors are input into a repeated word recognition network to obtain a matching judgment value; and multiple repeated word recognition networks with different parameters correspondingly obtain multiple matching judgment values.

[0114] Wherein, during the training process, the marked matching judgment value is used for training. If a sentence only contains repeated words, such as "I...I love you", "I" is used as an element in the first word vector, another "I" is used as an element in the second word vector, "love" is used as an element in the third word vector, and "you" is used as an element in the fourth word vector. The repeated word recognition network corresponding to the fully connected parameters of the first layer of neurons and the third layer of neurons and the fully connected parameters of the third layer of neurons and the fourth layer of neurons will be retained. The matching judgment value is set to 0. If a sentence is such as "harsh", the matching judgment value is set to 1.

[0115] In the detection process, the matching judgment value is used to find the corresponding self-vector that meets the content of deleting duplicate elements.

[0116] Among them, the word vectors are sequentially input into different layers of the repeated word recognition network according to the pronunciation time points from early to late. For example, the first pronunciation time point corresponding to the first column of the detection word matrix is ​​input into the first layer of the repeated word recognition network. For example, the second pronunciation time point corresponding to the second column of the detection word matrix is ​​input into the second layer of the repeated word recognition network.

[0117] Taking a value among multiple matching judgment values ​​that is greater than other matching judgment values ​​as a prediction matrix value;

[0118] The word vectors corresponding to the repeated word recognition network in the prediction matrix value are arranged in chronological order from early to late to obtain the original matching word matrix.

[0119] The element matching word matrix has time points as rows and matching words as columns.

[0120] Among them, if the input word vector is "I" as an element in the first word vector, another "I" as an element in the second word vector, "love" as an element in the third word vector, and "you" as an element in the fourth word vector. The repeated word recognition network corresponding to the fully connected parameters of the first layer of neurons and the third layer of neurons and the fully connected parameters of the third layer of neurons and the fourth layer of neurons will be retained. Then, according to the order of the first word vector, the second word vector and the third word vector, the original matching word matrix is ​​obtained.

[0121] Optionally, predicting missing words based on the speech Fourier image, multiple pronunciation sets and corresponding pronunciation time points to obtain a predicted word matrix and a detected word matrix includes:

[0122] According to the pronunciation set, find the character corresponding to each pronunciation to obtain a character set; obtain multiple character sets corresponding to multiple pronunciation time points;

[0123] According to the pronunciation time points from early to late, the corresponding word sets are arranged and filled into a two-dimensional matrix in sequence to obtain a detection word matrix; the rows of the detection word matrix represent the corresponding pronunciation time points; the columns of the detection word matrix represent the word set corresponding to a pronunciation time point;

[0124] The number of rows of the detection word matrix is ​​determined by the length of the speech signal. The number of columns of the detection word matrix is ​​equal to the number of elements of the word set with the largest number of elements.

[0125] Among them, as in this embodiment, the 1st second is taken as the pronunciation time point corresponding to the first pronunciation in the speech signal, and the last pronunciation of the speech signal corresponds to 4. The pronunciation time point of the 1st second is earlier than the pronunciation time point of the 2nd second, which is earlier than the pronunciation time point of the 3rd second, which is earlier than the pronunciation time point of the 4th second.

[0126] Based on the speech Fourier image and the detection word matrix, a blank word matrix and blank pronunciation time points are obtained; the elements of the columns corresponding to the blank pronunciations in the blank word matrix are 0.

[0127] The pronunciation time points adjacent to the blank pronunciation time points in the blank word matrix are used as predicted blank pronunciation time points.

[0128] Among them, the pronunciation time point that is earlier than the blank pronunciation time point and adjacent to the blank pronunciation time point, or the pronunciation time point that is later than the blank pronunciation time point and adjacent to the blank pronunciation time point in the blank word matrix is ​​used as the predicted blank pronunciation time point.

[0129] Perform word matching on the words corresponding to the predicted blank pronunciation time points in the blank word matrix to obtain a plurality of matching words;

[0130] The matching word includes a character in a predicted blank pronunciation time point and one or more characters in the database that can be matched.

[0131] According to the position of the characters in the matching word, a predicted blank character and a matching time point are obtained; the predicted blank character represents the characters in the matching word except the characters in the predicted blank pronunciation time point; the matching time point represents the blank pronunciation time point corresponding to the predicted blank character; one blank pronunciation time point corresponds to multiple predicted blank characters.

[0132] Among them, if all the columns adjacent to the pronunciation time point corresponding to the blank character matrix "我" are 0 in the previous column and all 0 in the next column, and the words obtained by matching are "我们" and "自我", then place "自" in the column adjacent to the pronunciation time point corresponding to "我" in the previous column, and place "们" in the column adjacent to the pronunciation time point corresponding to "我" in the next column.

[0133] Fill in the predicted blank characters corresponding to the matching time points in the blank character matrix to obtain the predicted character matrix.

[0134] Through the above method, predict the blanks corresponding to one pronunciation, and then, based on the original one, if it meets the requirements, do as follows

[0135] Optionally, obtaining the blank character matrix and the blank pronunciation time points based on the voice Fourier image and the detection character matrix includes:

[0136] Take the positions in the voice Fourier image where there is no user voice as blank pronunciations to obtain the blank pronunciation time points;

[0137] Among them, the method for obtaining the blank pronunciation time points is the same as the method for obtaining the pronunciation time points.

[0138] Among them, pass the image in the voice Fourier image through the target detection network, and according to the different frequencies of the user voice, find out which part is the position where there is no user voice;

[0139] Obtain a blank vector; the blank vector is a column vector with the same number of columns as the detection character matrix and all elements having a value of 0;

[0140] According to the time sequence of the blank pronunciation time points and the pronunciation time points corresponding to the detection character matrix, insert the blank vector into the detection character matrix to obtain the blank character matrix.

[0141] Among them, if the pronunciation time points corresponding to the detection character matrix are the 1st second, the 3rd second, and the 4th second, and the blank pronunciation time point is the 2nd second, then insert a blank vector to the right of the column with index 0 in the detection character matrix to obtain the blank character matrix.

[0142] Optionally, judging whether to add the predicted word according to the predicted character matrix and the original matching character matrix, and performing the prediction of the sentence meaning to obtain the predicted sentence includes:

[0143] Match the characters in multiple columns of the predicted character matrix and the characters in multiple columns of the original matching character matrix respectively to obtain multiple matching sentences.

[0144] Among them, perform the matching of words through the set keywords, word order, and context information.

[0145] The matching sentence is input into the temporal convolutional network, the fluency of the sentence is predicted, and the predicted sentence value is obtained.

[0146] Among them, the matching sentence is input into the event convolution network in the event convolution network in sequence from early to late according to the pronunciation time point corresponding to each word or the blank pronunciation time point.

[0147] In this embodiment, the temporal convolutional network is a temporal convolutional network (Temporal Convolutional Network, TCN).

[0148] Among them, one word corresponds to an output value of the temporal convolutional network.

[0149] Among them, when training the temporal convolutional network, the matching between the subject, predicate and object can be used to score it. The temporal convolutional network outputs a vector table containing 10 numbers, which is a score from 1 to 10. The predicted sentence value represents the score corresponding to the value in the vector that is greater than other probabilities.

[0150] The matching statement whose predicted statement value is greater than other predicted statement values ​​is taken as the predicted statement.

[0151] The matching sentence with the highest score is used as the predicted sentence.

[0152] Through the above method, by constructing a database containing multiple single matching words, the word with the most semantically satisfying meaning is found, and then the relationship between the verb and the attributive, the subject, the object and the predicate is scored to find the sentence with the most accurate recognition.

[0153] Optionally, based on the speech Fourier image, detecting corresponding pronunciations in the speech signal, finding similar pronunciations, and obtaining multiple pronunciation sets and corresponding pronunciation time points include:

[0154] Obtaining a plurality of constructed pronunciation sets; the constructed pronunciation sets comprising a plurality of similar pronunciations; the pronunciations representing the pronunciations corresponding to a character;

[0155] A pronunciation time window is obtained; the pronunciation time window represents the length corresponding to a pronunciation.

[0156] In this embodiment, since the horizontal axis of the speech Fourier image represents time and the vertical axis represents frequency, the pronunciation time window is set to a window of 1 second.

[0157] Detecting the area corresponding to the pronunciation time window in the speech Fourier image to obtain a detected pronunciation;

[0158] Finding a corresponding constructed pronunciation set as the pronunciation set according to the detected pronunciation;

[0159] Multiple pronunciation time windows correspond to obtaining multiple pronunciation sets;

[0160] The starting time point of the time period corresponding to the pronunciation time window is taken as the pronunciation time point.

[0161] In this embodiment, the time is counted from the beginning of a speech segment. If the pronunciation detected at the first second is "wo", the detected pronunciation is taken as "wo". In this example, "o", "huo" and "wo" are pronunciations in a constructed pronunciation set, so the constructed pronunciation set corresponding to "wo" is taken as the pronunciation set. The 1 at the first second is taken as the pronunciation time. At the same time, Beijing time, etc. can be used for counting.

[0162] Optionally, the sound playing device is used to play voice.

[0163] The algorithm and display provided herein are not inherently related to any particular computer, virtual system or other device. Various general purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing such systems. In addition, the present invention is not directed to any specific programming language either. It should be understood that various programming languages ​​can be utilized to realize the content of the present invention described herein, and the description of the above specific languages ​​is for disclosing the best mode of the present invention.

[0164] In the description provided herein, a large number of specific details are described. However, it is understood that embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.

[0165] Similarly, it should be understood that in order to streamline the present disclosure and aid in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the present invention, various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed method should not be interpreted as reflecting the intention that the claimed invention requires more features than those explicitly recited in each claim. More specifically, as reflected in the claims below, inventive aspects lie in less than all the features of the individual embodiments disclosed above. Therefore, the claims that follow the specific embodiment are hereby expressly incorporated into the specific embodiment, with each claim itself serving as a separate embodiment of the present invention.

[0166] Those skilled in the art will appreciate that the modules in the devices in the embodiments may be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments may be combined into one module or unit or component, and in addition they may be divided into a plurality of submodules or subunits or subcomponents. Except that at least some of such features and / or processes or units are mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this manner may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature providing the same, equivalent or similar purpose.

[0167] In addition, those skilled in the art will appreciate that, although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present invention and form different embodiments. For example, in the claims below, any one of the claimed embodiments may be used in any combination.

[0168] The various component embodiments of the present invention may be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or digital signal processor (DSP) may be used in practice to implement some or all of the functions of some or all of the components in the apparatus according to an embodiment of the present invention. The present invention may also be implemented as a device or device program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention may be stored on a computer-readable medium, or may have the form of one or more signals. Such a signal may be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0169] It should be noted that the above embodiments illustrate the present invention rather than limit it, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets shall not be construed as a limitation on the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware comprising a number of different elements and by means of a suitably programmed computer. In a unit claim enumerating a number of devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc., does not indicate any order. These words may be interpreted as names.

Claims

1. An interactive robot for hemiplegia and speech disorders, characterized in that: Including sound receiving device, sound playing device and processor; The processor is used in the following method: A speech signal and a speech Fourier image are obtained; the speech signal is speech with a language barrier received by a sound receiving device; the speech Fourier image represents a Fourier image converted from the speech signal; Based on the speech Fourier image, corresponding pronunciations in the speech signal are detected, similar pronunciations are found, and multiple pronunciation sets and corresponding pronunciation time points are obtained; one pronunciation time point corresponds to one pronunciation set; and the pronunciation set contains similar pronunciations; Based on the speech Fourier image, multiple pronunciation sets and corresponding pronunciation time points, predict missing words to obtain a prediction word matrix and a detection word matrix; Through a repeated word recognition network, based on multiple pronunciation sets and a detection word matrix, repeated words or words with the same pronunciation but different meanings in the speech signal are detected to obtain an original matching word matrix; The original matching word matrix represents a plurality of words constituting a speech; According to the predicted word matrix and the original matching word matrix, determining whether to add the predicted word, predicting the meaning of the sentence, and obtaining the predicted sentence; The predicted sentence is converted into a play voice signal and sent to a sound playing device.

2. The interactive robot for hemiplegia and speech disorders according to claim 1, characterized in that: The repeated word recognition network includes multiple layers of neurons; Each layer of neurons is controlled by a threshold structure; n layers of neurons correspond to n threshold structures; Each layer of neurons is fully connected to all other layers of neurons.

3. The interactive robot for hemiplegia and speech disorders according to claim 2, characterized in that: The repeated word recognition network detects repeated words or words with the same pronunciation but different meanings in the speech signal based on multiple pronunciation sets and the detection word matrix to obtain the original matching word matrix, including: The detection word matrix is ​​divided into columns to obtain a plurality of word vectors; one word vector corresponds to one pronunciation time point; an element in the word vector represents a word corresponding to the pronunciation corresponding to one pronunciation time point; According to the plurality of pronunciation sets, detecting whether pronunciation sets at adjacent pronunciation time points are the same pronunciation set; If the pronunciation sets of adjacent pronunciation time points are the same pronunciation set, the pronunciation time points corresponding to the corresponding word vectors are filled into the repeated pronunciation time set; Taking the word vectors corresponding to the pronunciation time points in the repeated pronunciation time set in the multiple word vectors as repeated word vectors to obtain a repeated word vector set; the elements in the repeated word vector set are arranged from early to late according to the pronunciation time points; Taking the word vectors corresponding to the pronunciation time points in the repeated pronunciation time set except for the word vectors in the multiple word vectors as non-repeated word vectors to obtain a non-repeated word vector set; the elements in the non-repeated word vector set are arranged from early to late according to the pronunciation time points; Based on the repeated pronunciation time set, the repeated word vector set and the non-repeated word vector set, an original matching word matrix is ​​obtained through a repeated word recognition network.

4. The interactive robot for hemiplegia and speech disorders according to claim 3, characterized in that: The method of obtaining an original matching word matrix based on the repeated pronunciation time set, the repeated word vector set and the non-repeated word vector set through a repeated word recognition network includes: The value in the threshold structure corresponding to all non-repeated word vectors in the non-repeated word vector set is set to 1; the value in the threshold structure being 1 indicates that only the parameters of neurons in the layer corresponding to the adjacent word vectors are retained for full connection; According to the number of pronunciation time points in the repeated pronunciation time set as the number of repetitions; Arrange the natural numbers between 0 and the sum of the number of repetitions plus 1 in descending order to obtain a set of repetition values; Using a repeated word vector in the repeated word vector set as a subscript repeated word vector; Taking the value in the set of repeated values ​​that is equal to the subscript of the subscript repeated word vector as the subscript repeated value; Put the subscript repetition value into the subscript repetition word vector corresponding threshold structure; Arrange the natural numbers between 0 and the sum of the value in the threshold structure plus 1 in descending order to obtain a set of connection layer numbers; the elements in the set of connection layer numbers represent the difference between the number of layers of neurons fully connected with the subscript repeated word vector and the number of layers of neurons corresponding to the subscript repeated word vector; The parameters of the full connection between the neuron corresponding to an element in the connection layer number set and the neuron corresponding to the subscript repeated word vector are retained; Multiple elements in the connection layer number set correspond to obtaining multiple repeated word recognition networks with different parameters; Based on the repeated word recognition networks with multiple different parameters and the multiple word vectors, an original matching word matrix is ​​obtained.

5. The interactive robot for hemiplegia and speech disorders according to claim 4, characterized in that: The repeated word recognition network based on the multiple different parameters and the multiple word vectors obtains an original matching word matrix, including: Inputting the plurality of word vectors into a repeated word recognition network to obtain a matching judgment value; a plurality of repeated word recognition networks with different parameters correspondingly obtain a plurality of matching judgment values; Taking a value among multiple matching judgment values ​​that is greater than other matching judgment values ​​as a prediction matrix value; The word vectors corresponding to the repeated word recognition network in the prediction matrix value are arranged in chronological order from early to late to obtain the original matching word matrix.

6. The interactive robot for hemiplegia and speech disorders according to claim 1, characterized in that: The method predicts missing words based on the speech Fourier image, multiple pronunciation sets and corresponding pronunciation time points to obtain a prediction word matrix and a detection word matrix, including: According to the pronunciation set, find the character corresponding to each pronunciation to obtain a character set; obtain multiple character sets corresponding to multiple pronunciation time points; According to the pronunciation time points from early to late, the corresponding word sets are arranged and filled into a two-dimensional matrix in sequence to obtain a detection word matrix; the rows of the detection word matrix represent the corresponding pronunciation time points; the columns of the detection word matrix represent the word set corresponding to a pronunciation time point; Based on the speech Fourier image and the detection word matrix, a blank word matrix and a blank pronunciation time point are obtained; the elements of the column corresponding to the blank pronunciation in the blank word matrix are 0; Taking the pronunciation time points adjacent to the blank pronunciation time points in the blank word matrix as predicted blank pronunciation time points; Perform word matching on the words corresponding to the predicted blank pronunciation time points in the blank word matrix to obtain a plurality of matching words; According to the position of the characters in the matching word, a predicted blank character and a matching time point are obtained; the predicted blank character represents the characters in the matching word except the characters in the predicted blank pronunciation time point; the matching time point represents the blank pronunciation time point corresponding to the predicted blank character; one blank pronunciation time point corresponds to multiple predicted blank characters; Fill the blank word matrix with the predicted blank words corresponding to the matching time point to obtain a predicted word matrix.

7. The interactive robot for hemiplegia and speech disorders according to claim 6, characterized in that: The method of obtaining a blank word matrix and a blank pronunciation time point based on the speech Fourier image and the detection word matrix comprises: The position where the user's voice does not exist in the voice Fourier image is regarded as a blank pronunciation to obtain a blank pronunciation time point; Obtain a blank vector; the blank vector is a column vector having the same number of columns as the detection word matrix and all element values ​​are 0; According to the time sequence of the blank pronunciation time points and the pronunciation time points corresponding to the detection word matrix, a blank vector is inserted into the detection word matrix to obtain a blank word matrix.

8. The interactive robot for hemiplegia and speech disorders according to claim 1, characterized in that: The step of determining whether to add a predicted word based on the predicted word matrix and the original matching word matrix, predicting the meaning of a sentence, and obtaining a predicted sentence includes: Matching the words in multiple columns in the predicted word matrix with the words in multiple columns in the original matching word matrix respectively to obtain multiple matching statements; Input the matching sentence into a temporal convolutional network, predict the fluency of the sentence, and obtain a predicted sentence value; The matching statement whose predicted statement value is greater than other predicted statement values ​​is taken as the predicted statement.

9. The interactive robot for hemiplegia and speech disorders according to claim 1, characterized in that: The method of detecting corresponding pronunciations in the speech signal based on the speech Fourier image, finding similar pronunciations, and obtaining multiple pronunciation sets and corresponding pronunciation time points includes: Obtaining a plurality of constructed pronunciation sets; the constructed pronunciation sets comprising a plurality of similar pronunciations; the pronunciations representing the pronunciations corresponding to a character; Obtaining a pronunciation time window; the pronunciation time window represents the length corresponding to a pronunciation; Detecting the area corresponding to the pronunciation time window in the speech Fourier image to obtain a detected pronunciation; Finding a corresponding constructed pronunciation set as the pronunciation set according to the detected pronunciation; Multiple pronunciation time windows correspond to obtaining multiple pronunciation sets; The starting time point of the time period corresponding to the pronunciation time window is taken as the pronunciation time point.

10. The interactive robot for hemiplegia and speech disorders according to claim 1, characterized in that: The sound playing device is used for playing voice.

Citation Information

Patent Citations

  • Chinese word pronunciation prediction method and device

    CN106910497A

  • Speaker recognition method based on convolution neural network and spectrogram

    CN106952649A