Human-computer interaction method and system for voice recognition

By building a user voice library and clustering algorithm, combining spectral map features and user confirmation coefficients, the problem of speech recognition technology adapting to individual differences is solved, and speech recognition and identity confirmation with high accuracy and security is achieved, improving the ease of use and reliability of human-computer interaction.

CN120279894AActive Publication Date: 2025-07-08广东公信智能会议股份有限公司

Patent Information

Application Number
CN202510620132.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-08
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

Existing speech recognition technology is difficult to adapt to the differences in individual language habits and speech characteristics, resulting in low identification accuracy and insufficient extraction of key information, affecting the security and reliability of human-computer interaction.

Method used

By building a user's voice library, using Fourier transform to generate a spectral map and clustering algorithm to divide word groups, combining upper and lower envelope features and user confirmation coefficients, dynamically adjusting the clustering results to achieve speech recognition and identity confirmation.

Benefits of technology

It improves the accuracy of speech recognition and the security of the system, enhances ease of use and response speed, adapts to the language habits and voice characteristics of different users, ensures the reliability of the recognition results and prevents illegal access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279894A_ABST
    Figure CN120279894A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech recognition, in particular to a man-machine interaction method and system for speech recognition, and the method comprises the steps: obtaining the frequency spectrum information of a user speech segment, generating a speech spectrogram, constructing a speech library of a user, and carrying out the word segmentation processing, thereby obtaining a sequence of word units; dividing word groups by using a clustering algorithm, and calculating a segmentation coefficient and a stop coefficient of each word group according to a mapping function of an upper envelope line and a lower envelope line of the spectrogram so as to determine an optimal clustering result; real-time input voice segments are compared with reference word groups, and a user confirmation coefficient between the spectrogram of the real-time word groups and the spectrogram of the reference word groups is calculated; and judging whether the real-time input voice segment and the reference voice segment are the same user or not according to the user confirmation coefficient, and completing the human-computer interaction process of voice recognition. According to the method, the user voice library is constructed, and the spectrogram analysis and clustering algorithm is utilized, so that efficient voice recognition and user confirmation are realized, and the user interaction experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition. In particular, it relates to a human - machine interaction method and system for speech recognition. Background Art

[0002] Speech recognition is a technology that converts human speech signals into text or instructions, enabling computers or other devices to understand and process human speech inputs. By analyzing the acoustic features of speech signals, a speech recognition system can identify words or commands in the speech and convert them into actionable information. The core of this technology lies in adapting to the language habits and speech characteristics of different users to achieve efficient human - machine interaction.

[0003] Speech recognition technology plays a crucial role in human - machine interaction. It not only improves the efficiency and convenience of interaction but also enhances the usability and intelligence level of devices. In fields such as smart homes, smart assistants, and in - vehicle systems, speech recognition enables users to communicate with devices through natural language without manual operation, thus achieving a seamless interaction experience. In addition, speech recognition provides a more user - friendly interaction method for disabled people, further promoting the popularization and application of the technology.

[0004] Although significant progress has been made in speech recognition technology, there are still some challenges in practical applications. First, when a speech recognition system processes the speech of different users, it is difficult to fully consider the unique language habits and speech characteristics of each person. For example, changes in sentence mood may cause significant differences in the pronunciation of the same word, affecting the accuracy of the system. Second, the identity recognition method based on a fixed speech template is difficult to adapt to these changes, reducing the security and reliability of human - machine interaction. In addition, existing technologies still have deficiencies in extracting and utilizing key information in speech, further limiting the accuracy and practicality of speech recognition. Summary of the Invention

[0005] To solve the problems of existing speech recognition technology, which is difficult to adapt to individual language habit and speech feature differences, resulting in low accuracy of identity recognition and insufficient extraction of key information, affecting the security and reliability of human - machine interaction, the present invention provides solutions in the following aspects.

[0006] In a first aspect, a human-computer interaction method for speech recognition includes: obtaining spectral information of a plurality of historical speech segments of a user, generating a spectrogram using Fourier transform, constructing a speech library of the user, performing word segmentation processing on each speech segment in the speech library to obtain a sequence of word units; using a clustering algorithm to divide the sequence of word units to obtain word groups composed of multiple consecutive word units, taking each word group in the speech library as a reference word group, and taking the spectrogram corresponding to each word group as a reference speech segment; the clustering process includes: calculating a segmentation coefficient for each word group according to a mapping function of the upper and lower envelope lines of the spectrogram, setting an initial number of clusters, distributing word units to the cluster clusters in order according to the segmentation coefficient, constructing a stopping coefficient for cluster cluster division according to the number of word units and the segmentation coefficient, and performing clustering iteration, and outputting an optimal clustering result based on the stopping coefficient; obtaining spectral information of a real-time input speech segment and the divided real-time word groups, and in response to the real-time word groups being the same as the reference word groups, retaining the spectrogram of the real-time word groups, and taking the similarity between the spectrogram of the real-time word groups and the spectrogram of the reference word groups as a user confirmation coefficient; judging whether the real-time input speech segment and the reference speech segment are from the same user according to the user confirmation coefficient, and completing the human-computer interaction process of speech recognition.

[0007] By constructing a user's speech library and reference word groups, generating spectrograms using Fourier transform, and dividing word groups using a clustering algorithm, the speech input of the user can be recognized more accurately. Adapting to the language habits and speech characteristics of different users, dynamically adjusting the clustering results by analyzing the upper and lower envelope lines and segmentation coefficients of the spectrogram to ensure the reliability of recognition. The user does not need to input a password or use other biometric methods, and only needs to input speech to complete identity confirmation, which simplifies the verification process, improves the usability of the system and user satisfaction. At the same time, by calculating the user confirmation coefficient, the system can judge whether the real-time input speech segment and the reference speech segment are from the same user, improving the security of the system and preventing illegal access. In addition, by optimizing the clustering process, over-segmentation is avoided, the performance of the system is improved, and the speech input of the user can be processed in real time, and the recognition result can be output quickly, enhancing the response speed of the system and the user's interaction experience.

[0008] Preferably, the performing word segmentation processing on each speech segment in the speech library to obtain a sequence of word units includes: Performing text conversion on each speech sentence in the speech library to obtain a corresponding text sentence, and using the jieba library to divide the text sentence according to word units to obtain a sequence of word units, where a text sentence contains several consecutive word units.

[0009] By converting the speech sentences in the speech library into text sentences and using the jieba library for word segmentation, the continuous text sentences can be divided into several word units, forming a sequence of word units. This not only improves the accuracy of speech recognition but also enhances the efficiency of text processing and the adaptability of the system, enabling better understanding and processing of various language structures and semantic information.

[0010] Preferably, the word unit is the basic component of speech or text and represents the smallest semantic unit of the speech signal or text content; the word group is a set composed of several consecutive word units and represents a specific part in speech or text; the text sentence is the complete text content after speech transcription and is a continuous text.

[0011] Preferably, the segmentation coefficient includes: Integrate the area between adjacent minimum points of the upper envelope, calculate the signal energy in the area between adjacent minimum points, divide the integration result by the product of the interval width and the maximum value between adjacent minimum points for normalizing the integration result, sum the normalized results of all adjacent minima and divide by the number of minimum points to obtain the mean value of the upper envelope; Integrate the area between adjacent maximum points of the lower envelope, calculate the signal energy in the area between adjacent maximum points, divide the integration result by the product of the interval width and the minimum value between adjacent maximum points for normalizing the integration result, sum the normalized results of all adjacent maxima and divide by the number of maximum points to obtain the mean value of the lower envelope; Take the average of the sum of the mean value of the upper envelope and the mean value of the lower envelope, and subtract the average value from 1 to obtain the segmentation coefficient of each word group.

[0012] By comprehensively analyzing the characteristics of the upper and lower envelopes of the speech signal, the high-frequency and low-frequency changes of the speech signal can be captured, helping the system to more accurately identify the subtle differences in speech, reduce the situation of misrecognition, adapt to the speech habits and emotional changes of different users, and ensure that the system can still provide reliable recognition results under variable speech inputs.

[0013] Preferably, the stopping coefficient includes: Calculate the ratio of the number of word units contained in each word group to the total number of word units in the whole text sentence to obtain the contribution degree of each word group; Take the average value of the sum of the products between the contribution degrees of all word groups and the segmentation coefficient as the stopping coefficient for the corresponding cluster division.

[0014] By comprehensively considering the contribution degree and feature significance of each word group, the quality of the clustering result can be evaluated more accurately, so as to determine the optimal number of clustering clusters, improve the accuracy and efficiency of clustering, and provide a more reliable basis for speech recognition and human-computer interaction.

[0015] Preferably, the output of the optimal clustering result based on the stop coefficient includes: When the stop coefficient is greater than the preset stop threshold or reaches the maximum value, the output clustering result is the optimal clustering result. Among them, in the optimal clustering result, each clustering cluster corresponds to a word group, and the number of the clustering clusters does not exceed the number of word groups.

[0016] By dynamically determining the optimal number of clustering clusters, over-segmentation is avoided, the accuracy and efficiency of the clustering result are ensured, and at the same time, the adaptability and robustness of the system are improved, providing reliable support for speech recognition and human-computer interaction.

[0017] Preferably, the user confirmation coefficient includes: Calculate the absolute difference between the spectrogram time length of the reference word group and the spectrogram time length of the real-time word group, and according to the ratio of the absolute difference to the reference time length, normalize the difference, and obtain the matching degree in terms of time length by subtracting the normalized ratio from 1, which is used to measure the similarity between the reference word group and the real-time word group in terms of time length; Use the DTW algorithm to calculate the similarity between the frequency sequence of the spectrogram of the reference word group and the frequency sequence of the spectrogram of the real-time word group, and take the ratio of the matching degree to the sum of the similarity between the frequency sequences and the hyperparameter as the user confirmation coefficient.

[0018] Preferably, the user confirmation coefficient further includes: Use the cosine similarity algorithm to calculate the similarity between the frequency sequences of the reference word group and the real-time word group; Calculate the absolute difference between the spectrogram time length of the reference word group and the spectrogram time length of the real-time word group, and according to the ratio of the absolute difference to the reference time length, normalize the difference, and obtain the matching degree in terms of time length by subtracting the normalized ratio from 1, which is used to measure the similarity between the reference word group and the real-time word group in terms of time length; Take the product of the similarity between the frequency sequences and the matching degree as the user confirmation coefficient.

[0019] Preferably, the determination of whether the real-time input speech segment and the reference speech segment are from the same user according to the user confirmation coefficient includes: If the user confirmation coefficient is less than or equal to the preset similarity threshold, it is considered that the input voice segment and the reference voice segment are not from the same user; otherwise, it is considered that the input voice segment and the reference voice segment are from the same user.

[0020] In a second aspect, a human-computer interaction system for speech recognition includes: a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned human-computer interaction method for speech recognition is implemented.

[0021] The present invention has the following effects: 1. By constructing a user's voice library and a reference word group, using Fourier transform to generate spectrograms and clustering algorithms to divide the word group, the present invention can capture the key features of voice signals and improve the accuracy of speech recognition. At the same time, by calculating the user confirmation coefficient, the system can more accurately determine whether the real-time input voice segment and the reference voice segment are from the same user, further improving the recognition accuracy.

[0022] 2. The present invention can adapt to the unique language habits and voice characteristics of different users. By analyzing the upper and lower envelopes of the spectrogram and the segmentation coefficient, the system can dynamically adjust the clustering results to ensure the reliability of the recognition results. Using the jieba library for word segmentation processing can better understand and process various language structures and semantic information, enhancing the adaptability of the system.

[0023] 3. By reducing the dependence on additional biometric technologies or complex password verification, the present invention can complete identity confirmation through voice input, which not only improves the usability of the system but also enhances the user's satisfaction with the system. At the same time, by calculating the user confirmation coefficient, the system can determine whether the real-time input voice segment and the reference voice segment are from the same user, improving the security of the system and preventing illegal access. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flowchart of the method from step S1 to step S4 in a human-computer interaction method for speech recognition according to an embodiment of the present invention.

[0025] Figure 2 is a structural block diagram of a human-computer interaction system for speech recognition according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.

[0027] Refer to Figure 1, A human-computer interaction method for speech recognition includes steps S1 - S4, specifically as follows: S1: Obtain the spectral information of several historical speech segments of the user, generate a spectrogram using Fourier transform, construct the user's speech library, perform word segmentation on each speech segment in the speech library, and obtain a sequence of word units.

[0028] Perform text conversion on each speech sentence in the speech library to obtain the corresponding text sentence, and use the jieba library to divide the text sentence according to word units to obtain a sequence of word units. Among them, a text sentence contains several consecutive word units.

[0029] A word unit is the basic component of speech or text and is used to represent the smallest semantic unit of a speech signal or text content; a word group is a set composed of several consecutive word units and is used to represent a specific part in speech or text; a text sentence is the complete text content after speech transcription and is a continuous text.

[0030] Exemplary: Taking a speech sentence: "Speech recognition technology is widely used in smart homes" as an example: The text sentence is: Speech recognition technology is widely used in smart homes.

[0031] The word groups are: Speech recognition technology, in smart homes, widely used.

[0032] The word units are: Speech, recognition, technology, in, smart, home, in, application, widely.

[0033] S2: Use a clustering algorithm to divide the sequence of word units to obtain word groups composed of multiple consecutive word units, use each word group in the speech library as a reference word group, and use the spectrogram corresponding to each word group as a reference speech segment.

[0034] Among them, the clustering process includes: According to the mapping function of the upper and lower envelopes of the spectrogram, calculate the segmentation coefficient of each word group, set the initial number of clusters, according to the segmentation coefficient, allocate the word units to the cluster clusters in order, according to the number of word units and the segmentation coefficient, construct a stopping coefficient for cluster cluster division, and perform clustering iteration, and output the best clustering result based on the stopping coefficient.

[0035] The segmentation coefficient includes: Integrate the area between adjacent minimum points of the upper envelope, calculate the signal energy in the area between adjacent minimum points, divide the integration result by the product of the interval width and the maximum value between adjacent minimum points to normalize the integration result, sum the normalized results of all adjacent minima and divide by the number of minimum points to obtain the mean value of the upper envelope; Integrate the region between adjacent maximum points of the lower envelope to calculate the signal energy within the region between adjacent maximum points. Divide the integration result by the product of the interval width and the minimum value between adjacent maximum points to normalize the integration result. Sum the normalized results of all adjacent maximum points and divide by the number of maximum points to obtain the mean value of the lower envelope; Take the average of the sum of the mean value of the upper envelope and the mean value of the lower envelope, and subtract the average from 1 to obtain the segmentation coefficient of each word or phrase group.

[0036] Specifically, the segmentation coefficient satisfies the following relational expression: ; In the formula, represents the segmentation coefficient of the th word or phrase group, represents the number of minimum values of the upper envelope, represents the th minimum value of the upper envelope, represents the th minimum value of the upper envelope, represents the mapping function of the upper envelope, represents the mapping function of the lower envelope, represents the abscissa corresponding to the th minimum value of the upper envelope, represents the maximum value between the th minimum value and the th minimum value of the upper envelope, represents the number of maximum values of the lower envelope, represents the th maximum value of the lower envelope, represents the th maximum value of the lower envelope, represents the minimum value between the th maximum value and the th maximum value of the upper envelope.

[0037] That is to say, reflects the high-amplitude region of the speech signal, reflects the low-amplitude region of the speech signal; By respectively taking the mean value of the ratio of the integral within two minimum value intervals in the upper envelope to the overall integral, the continuous change amplitude of the maximum values of the speech spectrogram can be expressed. When this value is small, it indicates that the continuous change amplitude in the high-amplitude region is large, the special degree of this section of speech in the high-amplitude region is good, and it is more likely to have discriminative features; By calculating the mean value of the ratio of the integrals of two maximum value intervals in the lower envelope to the overall integral respectively, the continuous change amplitude of the minimum values of the speech spectrogram can be expressed. When this value is small, it indicates that the continuous change amplitude in the low amplitude region is large, and the special degree of this section of speech in the low amplitude region is good, and it is more likely to have discriminative features.

[0038] The stopping coefficient includes: Calculate the ratio of the number of word units contained in each word group to the total number of word units in the overall text sentence to obtain the contribution degree of each word group; The average value obtained by summing the products of the contribution degrees of all word groups and the segmentation coefficient is used as the stopping coefficient for the corresponding cluster division.

[0039] Specifically, the stopping coefficient satisfies the following relational expression: ; In the formula, represents the stopping coefficient, represents the number of clusters, represents the th ratio of the number of word units contained in the th word group to the total number of word units in the overall text sentence, represents the segmentation coefficient of the

[0040] That is to say, the number of clusters is restricted by the proportion of the number of elements in the cluster (the number of word units in the word group) to the entire text sentence to prevent the growth of the value from directly equaling the number of word units in the text sentence when the stopping condition reaches the maximum value. Such clustering is meaningless, and the larger this proportion, the longer this word group, that is, it also has a better effect in distinguishing other text sentences in subsequent comparisons.

[0041] Based on the stopping coefficient, the best clustering result is output, including: When the stopping coefficient is greater than the preset stopping threshold or the stopping coefficient reaches the maximum value, the output clustering result is the best clustering result. Among them, in the best clustering result, each cluster corresponds to a word group, and the number of clusters does not exceed the number of word groups.

[0042] It should be noted that the preset stopping threshold is 0.6, which can be adjusted according to the specific implementation situation. When there are many word groups in the speech segment, the clustering process may generate a large number of clusters, which not only increases the computational complexity but also may lead to over-segmentation, making the features of each cluster not significant enough. In this case, using the stopping threshold can effectively control the clustering process and ensure the quality and efficiency of the clustering result.

[0043] S3: Obtain the spectral information of the real-time input voice segment and the divided real-time word groups. In response to the real-time word groups being the same as the reference word groups, retain the spectrogram of the real-time word groups, and use the similarity between the spectrogram of the real-time word groups and the spectrogram of the reference word groups as the user confirmation coefficient.

[0044] The user confirmation coefficient includes: Calculate the absolute difference between the spectrogram time length of the reference word groups and the spectrogram time length of the real-time word groups, and according to the ratio of the absolute difference to the reference time length, normalize the difference. Obtain the matching degree in terms of time length by taking the difference between 1 and the normalized ratio, which is used to measure the similarity between the reference word groups and the real-time word groups in terms of time length; Use the DTW algorithm to calculate the similarity between the frequency sequence of the spectrogram of the reference word groups and the frequency sequence of the spectrogram of the real-time word groups, and take the ratio of the matching degree to the sum of the similarity between the frequency sequences and the hyperparameters as the user confirmation coefficient.

[0045] Specifically, the user confirmation coefficient satisfies the following relational expression: ; In the formula, represents the user confirmation coefficient, represents the time length of the spectrogram of the reference word groups, represents the time length of the real-time spectrogram of the real-time word groups, represents the frequency sequence of the spectrogram of the reference word groups, represents the frequency sequence of the real-time spectrogram of the real-time word groups, represents the dynamic time warping algorithm.

[0046] That is to say, adding in the denominator is to avoid the situation where the denominator is zero. By comparing the time lengths of the reference word groups and the real-time word groups, the similarity of the voice signal in the time dimension is evaluated. The higher the similarity of the time length, the closer the rhythm and speech rate of the voice signal; by using the algorithm to compare the frequency sequences of the reference word groups and the real-time word groups, the similarity of the voice signal in the frequency dimension is evaluated. The higher the similarity of the frequency sequences, the closer the pitch and timbre of the voice signal. The larger the user confirmation coefficient, the more similar the characteristics of the input voice and the reference voice, and the higher the possibility of being confirmed as the same user.

[0047] In addition, another embodiment further includes: Use the cosine similarity algorithm to calculate the similarity between the frequency sequences of the reference word groups and the real-time word groups; Calculate the absolute difference between the spectrogram time length of the reference word group and the spectrogram time length of the real-time word group, and use the ratio of the absolute difference to the reference time length to normalize the difference. Obtain the matching degree in terms of time length by taking the difference between 1 and the normalized ratio, which is used to measure the similarity between the reference word group and the real-time word group in terms of time length; Take the product of the similarity between the frequency sequences and the matching degree as the user confirmation coefficient.

[0048] Specifically, the user confirmation coefficient satisfies the following relational expression: ; In the formula, represents the user confirmation coefficient, represents the time length of the spectrogram of the reference word group, represents the time length of the real-time spectrogram of the real-time word group, represents the cosine similarity algorithm, represents the frequency sequence of the spectrogram of the reference word group, represents the frequency sequence of the real-time spectrogram of the real-time word group.

[0049] That is to say, the calculation of cosine similarity is relatively simple and suitable for large-scale data sets. It cannot handle non-linear changes on the time axis, such as stretching or shifting of the time axis. When the frequency sequence lengths of the reference word group and the real-time word group are the same, cosine similarity can provide a fast and accurate similarity measure. For non-linear changes on the time axis, such as stretching or shifting of the time axis, it is suitable for processing speech signals with different speaking speeds. When the frequency sequence lengths of the reference word group and the real-time word group are different, the algorithm can find the optimal matching path through dynamic time warping. When there are obvious stretching or shifting of the two frequency sequences on the time axis, the algorithm can better align and compare the sequences.

[0050] S4: Judge whether the real-time input speech segment and the reference speech segment are from the same user according to the user confirmation coefficient, and complete the human-computer interaction process of speech recognition.

[0051] In response to the user confirmation coefficient being less than or equal to the preset similarity threshold, it is considered that the input speech segment and the reference speech segment are not from the same user. Otherwise, it is considered that the input speech segment and the reference speech segment are from the same user.

[0052] Exemplarily, the preset similarity threshold is 0.5, which can be adjusted according to the specific implementation situation.

[0053] The present invention also provides a human-computer interaction system for speech recognition. As Figure 2As shown, the system includes a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement a human-computer interaction method for speech recognition according to the first aspect of the present invention. The system also includes a communication bus, a communication interface, and other components well-known to those skilled in the art. Their settings and functions are known in the art and will not be elaborated herein.

[0054] It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent shall be subject to the appended claims.

Claims

1. A human-computer interaction method for speech recognition, characterized in that Including: Obtain the spectral information of several historical voice segments of the user, generate a spectrogram using Fourier transform, construct a voice library of the user, perform word segmentation on each voice segment in the voice library, and obtain a sequence of word units; Use a clustering algorithm to divide the sequence of word units, obtain word groups composed of multiple consecutive word units, use each word group in the voice library as a reference word group, and use the spectrogram corresponding to each word group as a reference voice segment; The clustering process includes: according to the mapping function of the upper and lower envelope lines of the spectrogram, calculate the segmentation coefficient of each word group, set the initial number of clusters, according to the segmentation coefficient, allocate the word units to the cluster clusters in order, according to the number of word units and the segmentation coefficient, construct a stopping coefficient for cluster cluster division, and perform clustering iteration, and output the best clustering result based on the stopping coefficient; Obtain the spectral information of the real-time input voice segment and the divided real-time word group. In response to the real-time word group being the same as the reference word group, retain the spectrogram of the real-time word group, and use the similarity between the spectrogram of the real-time word group and the spectrogram of the reference word group as the user confirmation coefficient; Judge whether the real-time input voice segment and the reference voice segment are of the same user according to the user confirmation coefficient, and complete the human-computer interaction process of voice recognition.

2. The human-computer interaction method for speech recognition according to claim 1, wherein The performing word segmentation on each voice segment in the voice library to obtain a sequence of word units includes: Perform text conversion on each voice sentence in the voice library to obtain a corresponding text sentence, and use the jieba library to divide the text sentence according to word units to obtain a sequence of word units, where a text sentence contains several consecutive word units.

3. A human-computer interaction method for speech recognition according to claim 2, characterized in that, The word unit is the basic component of voice or text, and is used to represent the smallest semantic unit of voice signals or text content; the word group is a set composed of several consecutive word units, and is used to represent a specific part in voice or text; the text sentence is the complete text content after voice transcription, and is a continuous text.

4. A human-computer interaction method for speech recognition according to claim 1, characterized in that The segmentation coefficient includes: integrate the region between adjacent minimum points of the upper envelope line, calculate the signal energy in the region between adjacent minimum points, divide the integration result by the product of the interval width and the maximum value between adjacent minimum points to normalize the integration result, sum the normalization results of all adjacent minimum values and divide by the number of minimum points to obtain the mean value of the upper envelope line; integrate the region between adjacent maximum points of the lower envelope line, calculate the signal energy in the region between adjacent maximum points, divide the integration result by the product of the interval width and the minimum value between adjacent maximum points to normalize the integration result, sum the normalization results of all adjacent maximum values and divide by the number of maximum points to obtain the mean value of the lower envelope line; take the average of the sum of the mean value of the upper envelope line and the mean value of the lower envelope line, and subtract the average value from 1 to obtain the segmentation coefficient of each word group.

5. A human-computer interaction method for speech recognition according to claim 1, characterized in that, The stopping coefficient includes: Calculate the ratio of the number of word units contained in each word group to the total number of word units in the overall text sentence to obtain the contribution degree of each word group; The average value obtained by summing the products of the contribution degrees of all word groups and the segmentation coefficient is used as the stopping coefficient for the corresponding clustering cluster division.

6. The human-computer interaction method for speech recognition according to claim 1, wherein Outputting the optimal clustering result based on the stopping coefficient includes: When the stopping coefficient is greater than the preset stopping threshold or reaches the maximum value, the output clustering result is the optimal clustering result. Among the optimal clustering results, each clustering cluster corresponds to a word group, and the number of clustering clusters does not exceed the number of word groups.

7. A human-computer interaction method for speech recognition according to claim 1, characterized in that, The user confirmation coefficient includes: Calculate the absolute difference between the spectrogram time length of the reference word group and the spectrogram time length of the real-time word group, and normalize the difference according to the ratio of the absolute difference to the reference time length. The difference between 1 and the normalized ratio is used to obtain the matching degree in terms of time length, which is used to measure the similarity between the reference word group and the real-time word group in terms of time length. Use the DTW algorithm to calculate the similarity between the frequency sequence of the spectrogram of the reference word group and the frequency sequence of the spectrogram of the real-time word group. The ratio of the sum of the matching degree and the similarity between the frequency sequences to the hyperparameter is used as the user confirmation coefficient.

8. A human-computer interaction method for speech recognition according to claim 1, characterized in that, The user confirmation coefficient further includes: Use the cosine similarity algorithm to calculate the similarity between the frequency sequence of the reference word group and the frequency sequence of the real-time word group. Calculate the absolute difference between the spectrogram time length of the reference word group and the spectrogram time length of the real-time word group, and normalize the difference according to the ratio of the absolute difference to the reference time length. The difference between 1 and the normalized ratio is used to obtain the matching degree in terms of time length, which is used to measure the similarity between the reference word group and the real-time word group in terms of time length. The product of the similarity between the frequency sequences and the matching degree is used as the user confirmation coefficient.

9. A human-computer interaction method for speech recognition according to claim 1, characterized in that, Judging whether the real-time input speech segment and the reference speech segment are from the same user according to the user confirmation coefficient includes: If the user confirmation coefficient is less than or equal to the preset similarity threshold, it is considered that the input speech segment and the reference speech segment are not from the same user; otherwise, it is considered that the input speech segment and the reference speech segment are from the same user.

10. A human-computer interaction system for speech recognition, characterized in that, It includes: A processor and a memory, where the memory stores computer program instructions. When the computer program instructions are executed by the processor, the human-computer interaction method for speech recognition according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Voiceprint identification method and apparatus

    CN106098068A

  • Speaker clustering method and a related device

    CN109800299A

  • Similarity data clustering method for dam safety monitoring data

    CN110197211A

  • Speaker recognition method based on word correlation score calculation

    CN110875044A

  • Speech processing apparatus, speech processing method, and non-transitory computer-readable medium storing program

    CN114175150A

Cited By

  • Personalized speech synthesis method, electronic equipment and system for digital human on AI platform

    CN120748362A

  • Method for personalized speech synthesis of AI platform digital human, electronic device and system

    CN120748362B

  • AI fused intelligent manufacturing workshop instruction processing output method and system

    CN122021579A