Intelligent voice recognition system and method based on internet of things
By combining the Internet of Things with an intelligent speech recognition system, and by using dynamic recognition step size and weighted algorithms to optimize speech templates, the problems of insufficient recognition range and accuracy in traditional systems have been solved, enabling wider and higher-precision speech recognition applications.
Patent Information
- Application Number
- CN202511288000.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-10
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-09-10
AI Technical Summary
Traditional intelligent voice recognition systems struggle to achieve flexible recognition across different scenarios and connected devices, resulting in insufficient voice recognition range and accuracy, thus limiting their application scenarios.
By using an IoT-based intelligent speech recognition system, which combines data import, recognition determination, analysis, and update modules, and utilizes dynamic recognition step size, feedback parameters, and weighted algorithms, speech templates are optimized to improve the accuracy and range of speech recognition.
It significantly expands the application scenarios of speech recognition, improves the accuracy and user experience of speech recognition, reduces latency, and protects privacy.
Smart Images

Figure CN120783732B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of Internet of Things intelligent voice recognition, in particular to an intelligent voice recognition system and method based on the Internet of Things. BACKGROUND
[0002] Intelligent voice recognition technology has penetrated into various aspects of our life and work by converting human voice into machine-readable instructions or text. In the field of smart home, voice recognition technology allows users to control various devices in the home through simple voice commands, improving the convenience and comfort of life.
[0003] Intelligent voice recognition is applied to various industries and life scenarios through intelligent voice recognition systems relying on language recognition large models. However, in traditional intelligent voice recognition systems, voice recognition mainly focuses on the voice feature recognition process of specific scenes, specific voice commands and language patterns, making it difficult to achieve flexible recognition of different scenes and connected device ends, resulting in limited application scenarios and insufficient voice recognition range and accuracy.
[0004] By combining the Internet of Things with voice recognition, integrating lightweight voice recognition models through some Internet of Things device ends and local gateways, and integrating a large number of sensors to directly understand some event environment data, voice commands can no longer control a single device, but can trigger a series of linkage scenario events, greatly expanding the use scenario range of voice recognition and significantly improving people's quality of life. SUMMARY
[0005] The purpose of the present application is to provide an intelligent voice recognition system and method based on the Internet of Things, which solves the following technical problems:
[0006] How to improve the voice recognition range and accuracy of users through a voice recognition model, optimize and update the voice recognition template.
[0007] The purpose of the present application can be achieved by the following technical solutions:
[0008] The intelligent voice recognition system based on the Internet of Things comprises:
[0009] A data import module is used to import pre-recorded voice templates from a database as first voice data, and to obtain a dynamic recognition step length of the first voice data as a first time. The dynamic recognition step length is the ratio of the recognition step length of a sentence in the voice data to the total number of word segmentation of the sentence, i.e. the recognition step length of a unit word segmentation in the sentence is obtained.
[0010] The recognition determination module is configured to input second voice data in real time, obtain a current dynamic recognition step of the second voice data as a second time, mark a difference between the first time and the second time as a feedback parameter, obtain all events related to the voice data in a voice template corresponding to the feedback parameter in a threshold interval, and mark the events as feedback events, and determine a feedback index value of the voice template according to the feedback events, and determine the actual recognition accuracy of the user voice based on the feedback index value.
[0011] The analysis module is configured to obtain part-of-speech features of a plurality of sentences based on the actual recognition accuracy of the user voice, perform online recognition correction, and obtain a correction result.
[0012] The update module is configured to substitute the correction result into the first voice data to obtain updated voice data of the voice template, and feed back the voice data to the database for updating.
[0013] Preferably, in the recognition determination module, the analysis manner of the feedback events includes:
[0014] obtaining a threshold interval corresponding to the feedback parameter, and screening a matched candidate event set from the voice template according to the threshold interval;
[0015] calculating a semantic correlation degree and a context fitting index of each event in the candidate event set and the current second voice data;
[0016] and generating an event correlation score based on a dynamic weighting algorithm to fuse the semantic correlation degree and the context fitting index;
[0017] selecting events with an event correlation score higher than a preset score threshold as the feedback events, and constructing the feedback event set.
[0018] Preferably, the feedback parameter further includes:
[0019] determining the size of the feedback parameter and the threshold interval, and comparing the feedback parameter with a positive threshold upper limit and a negative threshold lower limit of the threshold interval:
[0020] if the feedback parameter is greater than the positive threshold upper limit or less than the negative threshold lower limit, it is determined that an abnormality occurs, and a warning information is generated;
[0021] if the feedback parameter is between the negative threshold lower limit and the positive threshold upper limit, it is determined that it is normal, and the feedback events are obtained.
[0022] Preferably, the feedback index value is obtained in the following manner:
[0023] extracting historical recognition accuracy, environmental noise adaptation coefficient and user pronunciation stability parameters of each feedback event in the feedback event set;
[0024] The historical recognition accuracy, the environmental noise adaptation coefficient and the user pronunciation stability parameter are multi-dimensionally weighted and fused to generate a feedback index value. The calculation formula of the feedback index value is as follows:
[0025] ;
[0026] wherein, , , are weight coefficients, and , , are greater than 0, ; assuming that the feedback event set has events, each event contains the following parameters: historical recognition accuracy , environmental noise adaptation coefficient , and user pronunciation stability parameter ; and the values are all positive; is the minimum value in the historical recognition accuracy set of events, is the maximum value in the historical recognition accuracy set of events; is the minimum value in the environmental noise adaptation coefficient set of events, is the maximum value in the environmental noise adaptation coefficient set of events; is the minimum value in the user pronunciation stability parameter set of events, is the maximum value in the user pronunciation stability parameter set of events;
[0027] Based on the feedback index value , the actual recognition accuracy of the user voice is calculated in combination with the known transmission quality of the voice channel and the computing capability of the device end.
[0028] Preferably, the online recognition correction based on the actual recognition accuracy of the user voice in the analysis module is performed in the following manner to obtain a correction result:
[0029] It is judged whether the actual recognition accuracy of the user voice is less than a preset recognition accuracy threshold of the user voice:
[0030] If yes, the word segmentation sequences of multiple sentences of the user are corrected and analyzed to obtain the corrected word segmentation sequences.
[0031] If no, no correction and analysis are needed.
[0032] Preferably, the correction result further comprises:
[0033] After aligning the corrected word segmentation sequence with the corresponding sentence in the original first voice data, the similarity is calculated, and if the similarity is higher than a set threshold, the weighted smoothing method is used to integrate the correction result into the original voice template, iteratively update the voice template and feedback to the database.
[0034] Preferably, the similarity calculation method is:
[0035] The edit distance between the original word segmentation sequence and the corrected word segmentation sequence is calculated by alignment. ;
[0036] The edit distance is converted into a similarity score, normalized to the range , and the similarity is calculated as:
[0037] ;
[0038] wherein, and are the lengths of the sequences and , respectively.
[0039] The intelligent voice recognition method based on the Internet of Things, the method is run and realized by using an intelligent voice recognition system based on the Internet of Things, comprising:
[0040] S1, importing a pre-recorded voice template from a database as first voice data; and obtaining a dynamic recognition step of the first voice data as a first time; the dynamic recognition step is a ratio of a recognition step of a sentence in the voice data to a total number of word segmentation of the sentence, that is, the recognition step of a unit word segmentation in the sentence is obtained;
[0041] S2, real-time inputting second voice data, and obtaining a current dynamic recognition step of the second voice data as a second time, marking the difference between the first time and the second time as a feedback parameter; and obtaining all events related to the voice data in the voice template corresponding to the feedback parameter within a threshold interval, denoted as feedback events; determining a feedback index value of the voice template according to the feedback events and the feedback parameter, and determining the actual recognition accuracy of the user voice based on the feedback index value;
[0042] S3, based on the actual recognition accuracy of the user voice, obtaining the word segmentation features of multiple sentences for online recognition correction, and obtaining a correction result;
[0043] S4, substituting the correction result into the first voice data for fitting replacement, obtaining updated voice data of the voice template, and feeding back the voice data to the database for updating.
[0044] The beneficial effects of the present application: the data import module of the present application identifies the reference data information through the acquisition of the first time of the first voice data, and identifies the first time as the dynamic recognition step of the first voice data, the determination module identifies the control data, that is, the real-time input of the second voice data, and acquires the second time through the same recognition method, and marks the difference between the first time and the second time as a feedback parameter, sets a reasonable range for the feedback parameter, judges whether the confidence score is within this interval, then performs the next operation, further acquires all events related to voice data in the threshold interval, and records the feedback event; and determines the feedback index value of the voice template according to the feedback event, and identifies the accuracy of the user voice through the feedback index value, and further determines the actual recognition accuracy of the user voice based on the feedback index value.
[0045] Of course, implementing any product of the present application does not necessarily require all the advantages described above to be achieved at the same time. BRIEF DESCRIPTION OF DRAWINGS
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed for the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0047] Figure 1 The module diagram of the intelligent voice recognition system of the present application based on the Internet of Things;
[0048] Figure 2 The step diagram of the intelligent voice recognition method of the present application based on the Internet of Things. DETAILED DESCRIPTION
[0049] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0050] In the traditional intelligent speech recognition system, the speech feature recognition process mainly focuses on specific scenarios and language patterns. Unlike traditional intelligent speech recognition systems, this time, the Internet of Things and speech recognition are combined to solve some lightweight speech recognition models integrated on the Internet of Things device end and local gateway. There is no need to upload the cloud, and it can be directly processed on the local edge side, reducing the delay while protecting privacy. At the same time, the Internet of Things system can understand some environmental data of events through the integration of a large number of sensors, so that the voice command is no longer to control a single device, but can trigger some class of linkage scene events, greatly expanding the use scenario range of speech recognition and significantly improving people's quality of life. Therefore, the speech recognition range and recognition accuracy are obviously improved through the speech recognition system and method based on the Internet of Things.
[0051] Embodiment 1
[0052] Referring to Figure 1 The application is an intelligent speech recognition system based on the Internet of Things, which comprises:
[0053] A data import module is configured to import a pre-recorded voice template from a database as first voice data, and obtain a dynamic recognition step of the first voice data as a first time; the dynamic recognition step is a ratio of a recognition step of a sentence in the voice data to a total number of word segmentation of the sentence, that is, a recognition step of a unit word segmentation in the sentence is obtained;
[0054] A recognition determination module is configured to real-time input second voice data, and obtain a current dynamic recognition step of the second voice data as a second time, mark a difference between the first time and the second time as a feedback parameter, and obtain all events related to the voice data in the voice template corresponding to the feedback parameter in a threshold interval, denoted as feedback events; and determine a feedback index value of the voice template according to the feedback events, and determine an actual recognition accuracy of user voice based on the feedback index value;
[0055] An analysis module is configured to obtain word segmentation features of multiple sentences for online recognition correction based on the actual recognition accuracy of user voice, and obtain a correction result;
[0056] An update module is configured to substitute the correction result into the first voice data for fitting replacement to obtain voice data of an updated voice template, and feed back the voice data to the database for updating.
[0057] According to the above technical content, it should be pointed out that the data import module can import the pre-recorded voice template from the database as the first voice data as the reference data, and the voice template is usually a set of standard voice data pre-recorded, and the voice template also has a set of specific command words for starting the voice recognition operation; based on the specification information in the database for voice recognition, the first time of the first voice data is also acquired, the voice data is collected according to the current recognized voice information through the Internet of Things, and then the recognition step of the information of a word or a phrase (including prepositions) in a sentence is acquired; and the dynamic recognition step of the first voice data is acquired as the first time, and the dynamic recognition step is the ratio of the recognition step of the sentence in the voice data to the total number of word segmentation of the sentence, that is, the recognition step of the unit word segmentation in the sentence is acquired;
[0058] The recognition determination module acquires the dynamic recognition step of the current second voice data by acquiring the contrast data, that is, the real-time input second voice data, and acquiring the step in the same recognition manner, and takes the dynamic recognition step as the second time, and marks the difference between the first time and the second time as the feedback parameter as the quantitative performance index after voice interaction. The feedback parameter mainly reflects the confidence of the recognition engine on the recognition result. By setting a reasonable range for the feedback parameter, it is judged whether the confidence score is within this interval, and the next step is performed to further acquire all events related to voice data in the voice template corresponding to the threshold interval within the threshold interval, that is, all voice interaction records whose feedback parameters fall within the set threshold interval, which are recorded as feedback events. And determine the feedback index value of the voice template according to the feedback event, and identify the accuracy of the user's voice through the feedback index value, including the proportion of the number of feedback events to the total number of all voice events in the time period, reflecting the effectiveness of the voice recognition information under this voice template, and then determining the actual recognition accuracy of the user's voice based on the feedback index value.
[0059] The analysis module performs online recognition correction based on the actual recognition accuracy of the voice to realize the timely updating of the voice template by the updating module. Specifically, the online recognition correction is performed by acquiring the word segmentation features of multiple sentences, acquiring the correction result, and fitting and replacing the corrected first voice data to ensure that the updated voice template acquires the voice data, and realizes the timely updating and optimization process of the voice data in the database.
[0060] As an embodiment of the present application, the analysis method of the feedback event in the recognition determination module includes:
[0061] The threshold interval corresponding to the feedback parameter is acquired, and the matched candidate event set is screened out from the voice template according to the threshold interval;
[0062] calculate a semantic relevance degree and a context fitting index of each event in the candidate event set and the current second voice data;
[0063] generate an event relevance score based on dynamic weighting algorithm fusion of the semantic relevance degree and the context fitting index;
[0064] select events with an event relevance score higher than a preset score threshold as feedback events, and construct the feedback event set.
[0065] It needs to be explained that the analysis of the feedback events by the determination module is obtained by confidence interval analysis of the feedback parameters; first, by screening in the voice template within the threshold interval of the feedback parameters, the candidate event set in the difference range is quickly and preliminarily screened out from the huge voice template library (historical log); this step is mainly to reduce the subsequent calculation amount and exclude poor quality data; then, two key indicators are obtained, which are semantic relevance degree and context fitting index; the semantic relevance degree analyzes the similarity in meaning between the recognized text in the historical event and the recognized text of the current second voice data; advanced NLP models (such as BERT, Sentence-BERT, etc. word / sentence vector model) are used to convert the two texts into vectors in a high-dimensional space; the cosine similarity or Euclidean distance between the two vectors is calculated as the quantitative value of the semantic relevance degree, such as: the current instruction is "turn off the light in my room", and the historical instruction is "turn off the bedroom lighting"; the literal meaning is different, but the semantic relevance degree is very high; the context fitting index is used to analyze whether the context environment when the historical event occurs matches the context environment of the current instruction; such as: time context: whether they are both at night; whether they are both on weekends; location context: whether they occur in the same room; whether they are captured by the same device; user behavior sequence: whether the user has just performed a similar operation before issuing the instruction (such as turning on the music App and then asking the sound box to turn up the volume); device state: whether the device is in a similar state at the time (such as the TV being on); implementation usually requires constructing a feature vector to represent the context, and then calculating the fitting degree through a machine learning model or a rule-based algorithm;
[0066] Finally, a comprehensive event relevance score is calculated by weighting fusion method according to the specific scene; the dynamic adjustment of the weights of the two ensures that the proportion of semantic relevance and context fitting index in the scene can be selected according to different scene environments, improving the credibility of the event relevance score; and according to the time relevance score calculated by the weighting fusion method, the events with an event relevance score higher than a preset score threshold are selected as feedback events, and the feedback event set is constructed, ensuring that the selection is comprehensive in the dimensions of acoustics, semantics and context, and the data quality of the feedback event set is high.
[0067] As an embodiment of the present application, the feedback parameter further comprises:
[0068] The size of the feedback parameter and the threshold interval is judged, and the feedback parameter is compared with the upper limit of the positive threshold and the lower limit of the negative threshold of the threshold interval:
[0069] If the feedback parameter is greater than the upper limit of the positive threshold or less than the lower limit of the negative threshold, it is judged that an abnormality occurs, and a warning information is generated;
[0070] If the feedback parameter is between the lower limit of the negative threshold and the upper limit of the positive threshold, it is judged to be normal, and the feedback event is obtained.
[0071] It should be noted that the threshold interval is realized by taking the difference between the first time and the second time as the feedback parameter. The time difference has both positive and negative values. By assigning a reasonable threshold interval to the difference, the size is usually between the lower limit of the negative threshold (e.g., -0.3-0) and the upper limit of the positive threshold (e.g., 0-0.7). In this time difference interval , the change is reliable, so by comparing the size of the feedback parameter with the upper limit of the positive threshold 0.7 and the lower limit of the negative threshold -0.3 of the threshold interval, if the feedback parameter exceeds the interval, it indicates that an abnormality occurs or the equipment identified by the interval is abnormal, and a warning information needs to be generated in time; if it is within the interval, all events related to the voice data in the voice template corresponding to the threshold interval need to be obtained, that is, all voice interaction records whose feedback parameters fall within the set threshold interval are recorded as feedback events.
[0072] As an embodiment of the present application, the feedback index value is obtained in the following manner:
[0073] The historical recognition accuracy, environmental noise adaptation coefficient and user pronunciation stability parameters of each feedback event in the feedback event set are extracted;
[0074] The historical recognition accuracy, environmental noise adaptation coefficient and user pronunciation stability parameters are multi-dimensionally weighted and fused to generate a feedback index value. The calculation formula of the feedback index value is as follows:
[0075] ;
[0076] Among them, , , are weight coefficients, and , , are all greater than 0, ; assuming that the feedback event set has events, each event contains the following parameters: historical recognition accuracy , an environmental noise adaptation coefficient , a user pronunciation stability parameter ; and the values are all positive values; is the minimum value in the set of historical recognition accuracy rates of the event, is the maximum value in the set of historical recognition accuracy rates of the event; is the minimum value in the set of environmental noise adaptation coefficients of the event, is the maximum value in the set of environmental noise adaptation coefficients of the event; is the minimum value in the set of user pronunciation stability parameters of the event, is the maximum value in the set of user pronunciation stability parameters of the event;
[0077] Based on the feedback index value , in combination with the current known transmission quality of the voice channel and the computing capability of the device end, the actual recognition accuracy of the user voice is calculated.
[0078] It should be noted that the voice data of the voice interaction record is analyzed, the voice data is quantified through the Internet of Things intelligent recognition system, the historical recognition accuracy, the environmental noise adaptation coefficient and the user pronunciation stability parameter are obtained; then, the historical recognition accuracy, the environmental noise adaptation coefficient and the user pronunciation stability parameter are multi-dimensionally weighted and fused to generate a feedback index value, the feedback index value is calculated, the parameters corresponding to each event are normalized to obtain , , ; and weighted fusion processing is performed, that is, through the formula ; and the weight coefficient , , can be adjusted according to the field knowledge or experiment and other actual applications, in addition, if the maximum value of some parameter is equal to the minimum value (that is, the parameter is constant), the actual situation rarely appears, and the normalization denominator is zero, then it needs to be handled separately (such as directly setting the normalization value to 0 or using other normalization methods); the above parameters: historical recognition accuracy , environmental noise adaptation coefficient , user pronunciation stability parameter are used to ensure that the directions of all parameters are consistent, that is, the greater the value is, the better.
[0079] The feedback index value corresponding to the event is obtained, and the feedback index value is calculated based on the event The acquired feedback index value The accuracy of the user's speech recognition can be further calculated, combined with the current known transmission quality of the speech channel and the computing capability of the device end, to calculate the actual recognition accuracy of the user's speech, i.e., through the formula The recognition accuracy is calculated and obtained ; wherein, is a first preset weight coefficient, indicating the importance of the time accuracy in the overall accuracy; is a second preset weight coefficient, indicating the importance of the integrity in the overall accuracy; is a third preset weight coefficient, indicating the importance of the confidence in the overall accuracy; represents the time error, i.e., the deviation between the actual time and the predicted time (for example, in target tracking, the time difference between the detected target and the real time); is the absolute value of the time error, is the maximum allowed time error for normalization; represents the normalized time accuracy index, with a value range of [0, 1]; when the time error is 0, the item is 1; when the error approaches the maximum allowed error, the item approaches 0; represents the integrity index of the recognition system, such as the detection rate, coverage, or integrity score (such as IoU or recall rate in target detection), the specific meaning depends on the context, and the value range is usually [0, 1] or percentage representation; is the current event feedback index value, indicating the reliability of the event recognition result; represents the current event maximum possible confidence for normalization, is the normalized confidence, with a value range of [0, 1].
[0080] As an embodiment of the present application, the analysis module performs online recognition correction based on the actual recognition accuracy of the user's speech, and the way to obtain the correction result includes:
[0081] determining whether the actual recognition accuracy of the user's speech is less than the preset recognition accuracy threshold of the user's speech:
[0082] If yes, the word segmentation sequence of the user's multiple sentences is analyzed and corrected to obtain the corrected word segmentation sequence;
[0083] If no, no correction analysis is needed.
[0084] It should be noted that the correction of the actual recognition accuracy of the user voice adopts an online correction manner: a part-of-speech sequence of multiple sentences is extracted to construct a part-of-speech feature vector, including acoustic features, phoneme boundary stability and context confidence; the part-of-speech feature vector is input into a pre-trained recurrent neural network model to output part-of-speech error estimation and correction suggestions; the real-time recognition result is dynamically adjusted according to the correction suggestions to generate a corrected part-of-speech sequence as a correction result.
[0085] As an embodiment of the present application, the correction result further includes:
[0086] After aligning the corrected part-of-speech sequence with the corresponding sentence in the original first voice data, the similarity is calculated to determine whether the similarity is higher than a set threshold, and if so, the correction result is integrated into the original voice template by using a weighted smoothing method, the voice template is iteratively updated and fed back to the database.
[0087] It should be noted that the above machine learning process of the recognition model on the matched part-of-speech sequence, i.e., by calculating the alignment similarity of the corrected part-of-speech sequence, determining, correcting the part-of-speech sequence with a similarity higher than a set threshold, integrating the correction result into the original voice template by using a weighted smoothing method, realizing progressive update of the template, and iteratively optimizing the parameters of the recurrent neural network model according to the error feedback in the correction process to form a closed-loop learning mechanism; the updated voice template is returned to the database to complete online adaptive correction.
[0088] As an embodiment of the present application, the similarity calculation method is:
[0089] The edit distance between the original part-of-speech sequence and the corrected part-of-speech sequence is calculated by alignment. ;
[0090] The edit distance is converted into a similarity score and normalized to the range , and the similarity is calculated according to the formula:
[0091] ;
[0092] wherein, and are the lengths of the sequences and , respectively.
[0093] It should be noted that in the process of similarity analysis, the present application calculates the edit distance of the original part-of-speech sequence and the corrected part-of-speech sequence, and normalizes the edit distance between the two to the interval The greater the value, the greater the similarity, and the similarity is in the mode of 1 minus the ratio, wherein the denominator adopted is The difference between the two is that when based on the maximum length is more commonly used, it can avoid the similarity between the two from being excessively biased towards short sequences.
[0094] Embodiment 2
[0095] Please refer to Figure 2 As shown in the drawings, the intelligent voice recognition method based on the Internet of Things, the method is run and realized by using an intelligent voice recognition system based on the Internet of Things, comprising:
[0096] S1, import the pre-recorded voice template from the database as the first voice data; and obtain the dynamic recognition step length of the first voice data as the first time; the dynamic recognition step length is the ratio of the recognition step length of a sentence in the voice data to the total number of word segmentation of the sentence, that is, the recognition step length of the unit word segmentation in the sentence is obtained;
[0097] S2, real-time input the second voice data, and obtain the current dynamic recognition step length of the second voice data as the second time, mark the difference between the first time and the second time as the feedback parameter; and obtain all events related to the voice data in the voice template corresponding to the feedback parameter within the threshold interval, denoted as feedback events; determine the feedback index value of the voice template according to the feedback events and the feedback parameter, and determine the actual recognition accuracy of the user voice based on the feedback index value;
[0098] S3, based on the actual recognition accuracy of the user voice, obtain the word segmentation features of multiple sentences for online recognition correction, and obtain the correction result;
[0099] S4, substitute the correction result into the first voice data for fitting replacement, obtain the voice data of the updated voice template, and feed back the voice data to the database for updating.
[0100] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments. Especially, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.
[0101] The above is only an example and description of the concept of the present application. Those skilled in the art can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, as long as they do not deviate from the concept of the present application or exceed the scope defined by the present application.
Claims
1. An intelligent voice recognition system based on the Internet of Things, characterized in that, include: The data import module is used to import pre-recorded voice templates from the database as the first voice data; And obtain the dynamic recognition step size of the first voice data as the first moment; The dynamic recognition step size is the ratio of the recognition step size of a sentence in the speech data to the total number of words in the sentence, that is, to obtain the recognition step size of a unit word in the sentence. The recognition and determination module is used to record the second voice data in real time, obtain the current dynamic recognition step size of the second voice data as the second moment, and mark the difference between the first moment and the second moment as a feedback parameter. And obtain all events related to voice data within the voice template corresponding to the feedback parameters within the threshold range, and record them as feedback events; The feedback index value of the voice template is determined based on the feedback event, and the actual recognition accuracy of the user's voice is determined based on the feedback index value. The analysis module is used to obtain the word segmentation features of multiple sentences based on the actual recognition accuracy of the user's speech, perform online recognition correction, and obtain the correction results; The update module is used to substitute the correction result into the first speech data for fitting and replacement, obtain the speech data of the updated speech template, and feed the speech data back to the database for updating. The analysis methods for feedback events in the identification and determination module include: Obtain the threshold range corresponding to the feedback parameters, and filter out the matching candidate event set from the speech template based on the threshold range; Calculate the semantic relevance and contextual fit index of each event in the candidate event set with the current second speech data; It also generates an event relevance score by fusing semantic relevance and context fit index based on a dynamic weighting algorithm; Events with relevance scores higher than a preset score threshold are selected as feedback events, and this set of feedback events is constructed. The method for obtaining the feedback indicator value is as follows: Extract the historical recognition accuracy, environmental noise adaptation coefficient, and user pronunciation stability parameters of each feedback event in the feedback event set; The historical recognition accuracy, environmental noise adaptation coefficient, and user pronunciation stability parameters are weighted and fused from multiple dimensions to generate a feedback index value; the calculation formula for the feedback index value is as follows: ; in, , , All are weighting coefficients, and , , All are greater than 0. Assume the set of feedback events has One event, Each event Includes the following parameters: historical recognition accuracy Environmental noise adaptability factor User pronunciation stability parameters And all values are positive. for The minimum value among the historical recognition accuracy sets of a given event. for The maximum value among the historical recognition accuracy sets of each event; for The minimum value among the set of environmental noise fitness coefficients for each event. for The maximum value in the set of environmental noise adaptation coefficients for each event; for The minimum value in the set of user pronunciation stability parameters for each event. for The maximum value in the set of user pronunciation stability parameters for each event; Based on this feedback indicator value Based on the known transmission quality of the voice channel and the computing power of the device, the actual recognition accuracy of the user's voice is calculated.
2. The IoT-based intelligent voice recognition system according to claim 1, characterized in that, The feedback parameters also include: Determine the magnitude of the feedback parameter relative to the threshold range, and compare the feedback parameter with the upper positive threshold and the lower negative threshold of the threshold range: If the feedback parameter is greater than the upper limit of the positive threshold or less than the lower limit of the negative threshold, it is determined that an anomaly has occurred and an early warning message is generated. If the feedback parameter is between the lower limit of the negative threshold and the upper limit of the positive threshold, it is judged as normal, and a feedback event is obtained.
3. The intelligent voice recognition system based on the Internet of Things according to claim 1, characterized in that, The analysis module performs online recognition correction based on the actual recognition accuracy of the user's speech, and the correction results are obtained in the following ways: Determine whether the actual recognition accuracy of the user's voice is less than the preset recognition accuracy threshold for the user's voice: If so, then perform correction analysis on the word segmentation sequences of multiple sentences from the user to obtain the corrected word segmentation sequences; If not, then no further analysis is needed.
4. The IoT-based intelligent voice recognition system according to claim 3, characterized in that, The correction results also include: The similarity is calculated after aligning the corrected word segmentation sequence with the corresponding sentence in the original first speech data. If the similarity is higher than a set threshold, a weighted smoothing method is used to integrate the correction result into the original speech template, iteratively update the speech template, and feed it back to the database.
5. The IoT-based intelligent voice recognition system according to claim 4, characterized in that, The similarity calculation method is as follows: The original word segmentation sequence is calculated by alignment. and the corrected word segmentation sequence Edit distance between ; Convert edit distance to similarity score and normalize to range. Internal similarity The calculation formula is: ; in, and Sequences and The length.
6. An intelligent speech recognition method based on the Internet of Things, characterized in that, The method employs the IoT-based intelligent voice recognition system as described in any one of claims 1-5, and the method includes: S1. Import a pre-recorded speech template from the database as the first speech data; and obtain the dynamic recognition step size of the first speech data as the first moment; the dynamic recognition step size is the ratio of the recognition step size of the sentence in the speech data to the total number of words in the sentence, that is, obtain the recognition step size of the unit word in the sentence. S2. Real-time recording of the second voice data, and obtaining the current dynamic recognition step size of the second voice data as the second moment, marking the difference between the first moment and the second moment as a feedback parameter; and obtaining all events related to the voice data within the voice template corresponding to the feedback parameter within the threshold range, denoted as feedback events; determining the feedback index value of the voice template based on the feedback events and feedback parameters, and determining the actual recognition accuracy of the user's voice based on the feedback index value; the analysis method of the feedback events includes: Obtain the threshold range corresponding to the feedback parameters, and filter out the matching candidate event set from the speech template based on the threshold range; Calculate the semantic relevance and contextual fit index of each event in the candidate event set with the current second speech data; It also generates an event relevance score by fusing semantic relevance and context fit index based on a dynamic weighting algorithm; Events with relevance scores higher than a preset score threshold are selected as feedback events, and this set of feedback events is constructed. The method for obtaining the feedback indicator value is as follows: Extract the historical recognition accuracy, environmental noise adaptation coefficient, and user pronunciation stability parameters of each feedback event in the feedback event set; The historical recognition accuracy, environmental noise adaptation coefficient, and user pronunciation stability parameters are weighted and fused from multiple dimensions to generate a feedback index value; the calculation formula for the feedback index value is as follows: ; in, , , All are weighting coefficients, and , , All are greater than 0. Assume the set of feedback events has One event, Each event Includes the following parameters: historical recognition accuracy Environmental noise adaptability factor User pronunciation stability parameters And all values are positive. for The minimum value among the historical recognition accuracy sets of a given event. for The maximum value among the historical recognition accuracy sets of each event; for The minimum value among the set of environmental noise fitness coefficients for each event. for The maximum value in the set of environmental noise adaptation coefficients for each event; for The minimum value in the set of user pronunciation stability parameters for each event. for The maximum value in the set of user pronunciation stability parameters for each event; Based on this feedback indicator value Based on the known transmission quality of the voice channel and the computing power of the device, the actual recognition accuracy of the user's voice is calculated. S3. Based on the actual recognition accuracy of the user's speech, obtain the word segmentation features of multiple sentences for online recognition and correction, and obtain the correction results; S4. Substitute the correction result into the first speech data for fitting and replacement to obtain the speech data of the updated speech template, and feed the speech data back to the database for updating.
Citation Information
Patent Citations
Remote voice recognition system and method with degree of association
CN106971722A
Smart home voice recognition control system based on Internet of Things
CN111768778A