Humanoid robot control method and system based on auditory sense
By analyzing noise and position information for denoising, segmenting and correlating voice signals, and generating control instructions, the problem of robots being difficult to recognize long voices is solved, and the accuracy and effect of robot control is improved.
Patent Information
- Application Number
- CN202510192143.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
AI Technical Summary
Existing robot systems are difficult to effectively recognize and process long voice commands, resulting in low control accuracy and effect.
By collecting and transmitting voice signals, analyzing noise conditions and transmitting position information for denoising, extracting voice signal characteristics and processing them in segments, converting them into text content segments, analyzing the correlation degree to generate control instructions.
It improves the robot's ability to recognize long speech, enhances the reliability of voice signals, the accuracy and effect of robot control, and ensures the stability of command execution.
Smart Images

Figure CN120048259A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech analysis, and particularly to an auditory-based humanoid robot control method and system. Background Art
[0002] In the development of robot control technology, auditory-based humanoid robot control solutions are gradually becoming a research hotspot. However, most current robot systems face a significant technical limitation: they can only recognize short voice commands and have relatively weak recognition capabilities for long voices.
[0003] For long voice commands, especially those containing complex semantic structures and multiple task requirements, existing robots often have difficulty effectively recognizing and processing them. This is mainly because long voice commands involve more vocabulary, grammar structures, and context information, posing higher requirements for the robot's speech recognition and semantic understanding capabilities.
[0004] In the prior art, robots often cannot recognize long voice commands, resulting in low precision and effectiveness of robot control and unable to ensure that the robot quickly and accurately recognizes commands.
[0005] Therefore, how to recognize long voices to improve the precision and effectiveness of robot control is a technical problem to be solved currently. Summary of the Invention
[0006] The purpose of the present invention is to solve the problem in the prior art that due to poor long voice recognition ability, the precision and effectiveness of robot control are low, and an auditory-based humanoid robot control method is proposed, which includes: Collect the voice signal of the person giving the command, analyze the noise situation of the voice signal and the position information of the person giving the command, and perform noise reduction processing on the voice signal according to the noise situation and the position information of the person giving the command; Extract multiple features of the voice signal, segment the voice signal to obtain multiple voice signal segments; Convert the multiple voice signal segments into multiple text content segments, analyze the correlation degree of the text content segments, and correlate the multiple text content segments to obtain semantic content; Extract multiple action behaviors from the semantic content, integrate the action behaviors to generate a control command, and manipulate the humanoid robot to perform corresponding activities through the control command.
[0007] In some embodiments of the present application, analyzing the noise situation of the voice signal and the position information of the person giving the command includes: Obtain the voice signal data containing noise in the past, extract modeling features, determine the M value based on the AIC rule and the BIC rule, initialize the parameters of the Gaussian component, and iteratively optimize the parameters of the Gaussian component to train the GMM model; The noise of the speech signal is recognized through the GMM model to obtain the spectrogram of the noise, the noise situation is confirmed on the spectrogram of the noise, and the position information of the speaker is determined through the sound source localization technology; Among them, the noise situation includes the spectral range width of the noise, the spectral peak change rate, and the energy distribution uniformity.
[0008] In some embodiments of the present application, the speech signal is denoised according to the noise situation and the position information of the speaker, including, Based on the spectral range width of the noise, the order of the filter is set, based on the spectral peak change rate and the energy distribution uniformity, the step size of the filter is set, the noise is denoised through the order and step size of the filter, and the speech signal at the position information of the speaker is enhanced.
[0009] In some embodiments of the present application, multiple features of the speech signal are extracted, including, The highest amplitude points at different frequencies in the spectrogram of the speech signal are connected to form a spectral envelope line, and smoothing processing is performed to obtain a spectral envelope diagram; A number of marker points are evenly set on the spectral envelope diagram, and the peak points and valley points are marked. The characteristic parameters are calculated by analyzing the marker points, peak points, and valley points on the spectral envelope diagram. The characteristic parameters include peak difference, valley difference, slope, and entropy value. The peak difference is the amplitude difference between the marker point and its nearest peak point, and the valley difference is the amplitude difference between the marker point and its nearest valley point.
[0010] In some embodiments of the present application, the speech signal is segmented to obtain multiple speech signal segments, including, Based on the peak difference, valley difference, slope, and entropy value, the fluctuation index of the marker points on the spectral envelope diagram is calculated, and the speech signal is segmented through the fluctuation index to obtain multiple speech signal segments; ; Among them, is the fluctuation index of the th marker point, is the fluctuation conversion coefficient, , are respectively the peak difference and valley difference of the th marker point, represents the minimum value of the peak difference and valley difference of the th marker point, , are respectively the combined conversion coefficients of the slope and entropy value, is the slope of the th marker point, is a preset constant.
[0011] In some embodiments of the present application, multiple speech signal segments are converted into multiple text content segments, and the relevance of the text content segments is analyzed, including: Each speech signal segment is input into a preset acoustic model to output the corresponding text content segment of each speech signal segment, and redundant content is removed from each text content segment, where the redundant content includes HTML tags and special characters; Each text content segment is subjected to word segmentation, and the part-of-speech of each word after word segmentation is labeled. The text content segment is parsed by a syntactic analyzer to identify the components in the text content segment and the relationships between the components; The part-of-speech coverage and the tightness of the component relationships of each text content segment are analyzed, and the integrity and missing information of the text content segment are defined in combination with the part-of-speech coverage and the tightness of the component relationships; When the integrity of the text content segment exceeds a preset integrity threshold, the text content segment is used as a complete text segment; otherwise, the text content segment is used as an incomplete text segment; For an incomplete text segment, the missing information is used to check whether the text content segments in the context of the incomplete text segment match, and the incomplete text segment is spliced with the text content segments in the context to be converted into a complete text segment; The relevance is defined by analyzing the complete text segments and the incomplete text segments.
[0012] In some embodiments of the present application, the relevance is defined by analyzing the complete text segments, including: The key features in the complete text segments and the incomplete text segments are extracted, each key feature is quantified and standardized, the respective sentence feature sets of each complete text segment and each incomplete text segment are constructed, and the relevance is calculated; ; wherein, is the relevance between the th text segment and the th text segment, is the sentence feature set of the th text segment, is the sentence feature set of the th text segment, represents the feature with the greatest contribution in the intersection of the sentence feature sets of the th text segment and the th text segment, is an adjustment factor, represents the adjustment factor obtained by mapping the feature with the greatest contribution in the intersection of the sentence feature sets, 、 respectively represent the weights of the th text segment and the th text segment, , respectively represent the integrity of the th text segment and the th text segment, where is a preset constant.
[0013] In some embodiments of the present application, multiple text contents are associated to obtain semantic content, including classifying multiple text content segments according to the association degree between pairwise text content segments, extracting the core semantic content of each category, and obtaining semantic content.
[0014] Correspondingly, the present application further provides an auditory-based humanoid robot control system, including a denoising module for collecting the voice signal of the order giver, analyzing the noise situation of the voice signal and the position information of the order giver, and performing denoising processing on the voice signal according to the noise situation and the position information of the order giver; a segmentation module for extracting multiple features of the voice signal, segmenting the voice signal to obtain multiple voice signal segments; an association module for converting multiple voice signal segments into multiple text content segments, analyzing the association degree of the text content segments, and associating multiple text content segments to obtain semantic content; a control module for extracting multiple action behaviors from the semantic content, integrating the action behaviors to generate a control instruction, and manipulating the humanoid robot to perform corresponding activities through the control instruction.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. Analyze the noise situation of the voice signal and the position information of the order giver, perform denoising processing on the voice signal according to the noise situation and the position information of the order giver, enhance the voice signal specifically, and effectively remove noise, improving the reliability of the voice signal data and providing a basis for subsequent voice analysis and recognition.
[0016] 2. Segment the voice signal to obtain multiple voice signal segments, analyze the association degree of the text content segments, and associate multiple text content segments to obtain semantic content. Segment the voice according to the features of the voice signal, and perform text content association according to the association situation of the text content segments to obtain accurate semantics, thereby improving the robot's recognition ability for long voices. The accuracy and effect of robot control are improved, and the stability of robot command execution is ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a schematic flow chart of an auditory-based humanoid robot control method proposed by the present invention; Figure 2Schematic diagram of the structure of an auditory-based humanoid robot control system proposed by the present invention. Detailed implementation manners
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0019] Refer to Figure 1 , an auditory-based humanoid robot control method, including the following steps, Step S101, collect the voice signal of the issuer, analyze the noise situation of the voice signal and the position information of the issuer, and perform noise reduction processing on the voice signal according to the noise situation and the position information of the issuer.
[0020] In this embodiment, in the field of human-computer interaction, controlling a humanoid robot to perform tasks through voice signals is one of the key technologies to improve the intelligence level of the robot. However, existing robot voice control systems are mostly limited to recognizing and processing short and clear voice commands, and have limited ability to process long voice commands in complex environments. For this reason, we propose an auditory-based intelligent control scheme for humanoid robots, aiming to improve the robot's understanding and execution ability of complex voice signals.
[0021] In this embodiment, the voice signal of the issuer is collected through a high-precision microphone array (set on both sides of the head of the humanoid robot), and then the collected signal is subjected to noise analysis and issuer position positioning. An advanced noise reduction algorithm is used to preprocess the voice signal to eliminate the interference of environmental noise, and at the same time, sound source localization technology is used to enhance the clarity of the issuer's voice signal.
[0022] In some embodiments of the present application, analyzing the noise situation of the voice signal and the position information of the issuer includes, Obtain the voice signal data containing noise in the past, extract the modeling features, determine the M value based on the AIC rule and the BIC rule, initialize the parameters of the Gaussian components, and iteratively optimize the parameters of the Gaussian components to train the GMM model; Perform noise recognition on the voice signal through the GMM model to obtain the spectrogram of the noise, confirm the noise situation on the spectrogram of the noise, and determine the position information of the issuer through sound source localization technology; Among them, the noise situation includes the spectral range width of the noise, the spectral peak change rate, and the energy distribution uniformity.
[0023] In this embodiment, models of speech signals and noise are established, such as Hidden Markov Model (HMM), Gaussian Mixture Model (GMM), etc. The speech signal to be processed is matched with these models, and whether it is noise is judged according to the matching degree. A large number of speech signal and noise samples are collected. These samples should cover various possible speech and noise types to ensure the generalization ability of the model. Feature extraction is performed on the collected speech signal and noise samples. Commonly used features include Mel Frequency Cepstral Coefficients (MFCC), spectral features, time-domain features, etc. The purpose of feature extraction is to convert the original signal into a numerical representation that can reflect its essential characteristics, facilitating subsequent modeling and matching. For the Gaussian Mixture Model (GMM), for GMM, parameters such as the number of Gaussian components, mean, and variance need to be determined, and these parameters are estimated through optimization methods such as the Expectation-Maximization (EM) algorithm. The number of Gaussian components (M) is an important parameter of the GMM model, which determines the fitting ability and complexity of the model to the data. Based on the model selection criterion: Use model selection criteria (such as AIC, BIC) to evaluate the performance of the model under different M values, and select the M value that makes the criterion reach the optimal. AIC criterion (Akaike Information Criterion): AIC is a standard for measuring the goodness of fit of a statistical model. It takes into account the complexity of the model and the quality of the data fitting. The smaller the AIC value, the better the model. When selecting the M value, multiple GMM models with different M values can be trained, and the AIC value of each model is calculated. Select the M value corresponding to the model with the smallest AIC value as the optimal value. BIC criterion (Bayesian Information Criterion): BIC is similar to AIC, but BIC punishes the model complexity more severely, so it tends to select a model with fewer parameters. Similarly, the optimal M value can be determined by calculating the BIC values of the GMM models under different M values. Combine the AIC rule and the BIC rule to jointly determine the M value. The parameters of the Gaussian components include the mean and variance, etc.
[0024] In this embodiment, the process of iteratively optimizing the parameters of the Gaussian components is as follows: When optimizing the parameters of the Gaussian Mixture Model (GMM) through the Expectation-Maximization (EM) algorithm, it mainly involves two iterative steps: the E step (Expectation step) and the M step (Maximization step). These two steps are executed alternately until the parameter estimation converges to a local optimal solution.
[0025] E step (Expectation step): Calculate the expected value of the hidden variable: In GMM, the latent variable usually refers to the probability that each data point belongs to each Gaussian component. In the E-step, we need to calculate the posterior probability that each data point belongs to each Gaussian component based on the current parameter estimates (mean, variance, and weight).
[0026] Using Bayes' theorem: The posterior probability can be calculated by Bayes' theorem, that is, by using the current data points and parameter estimates to update the probability that each data point belongs to each Gaussian component.
[0027] M-step (maximization step): Maximize the likelihood function: In the M-step, we use the posterior probability calculated in the E-step to update the parameters (mean, variance, and weight) of the GMM to maximize the likelihood function of the data. This usually involves taking the partial derivative of the likelihood function and setting the partial derivative to zero to solve for the optimal parameters.
[0028] Update parameters: For the mean, the update formula is: the mean of each Gaussian component is equal to the weighted average of all data points under that component, and the weight is the posterior probability that the data point belongs to that component.
[0029] For the variance, the update formula is: the variance of each Gaussian component is equal to the weighted average of the squares of the differences between all data points and the mean under that component, and the weight is also the posterior probability.
[0030] For the weight, the update formula is: the weight of each Gaussian component is equal to the sum of the posterior probabilities of all data points under that component.
[0031] Iteration and convergence: Repeat the E-step and M-step: The E-step and M-step are executed alternately until the change in parameter estimates is less than a preset threshold or the preset number of iterations is reached. This usually means that the algorithm has converged to a local optimal solution.
[0032] Convergence check: After each iteration, the change in parameter estimates can be checked. If the change is less than a certain threshold, the algorithm can be considered to have converged.
[0033] Use the trained GMM model to perform noise recognition on new speech signals. For each speech signal frame, calculate the posterior probability that it belongs to the noise Gaussian component. If the posterior probability of a certain frame exceeds a certain threshold, then that frame is considered to contain noise.
[0034] In this embodiment, the sound source localization technology refers to measuring sound signals at different positions in the environment using multiple microphones, and determining the position of the sound source by analyzing parameters such as the time difference, phase difference, or frequency difference of the sound signals. The signals received by the microphone array are subjected to time alignment and synchronization processing. Then, methods such as cross-correlation analysis and generalized cross-correlation are used to estimate the time difference of the sound signals arriving at different microphones. Finally, based on the time difference and the geometric layout of the microphones, algorithms such as triangulation or hyperbolic positioning are used to calculate the position of the sound source.
[0035] In some embodiments of the present application, the speech signal is denoised according to the noise situation and the position information of the speaker, including, The order of the filter is set based on the spectral range width of the noise, the step size of the filter is set based on the spectral peak change rate and the energy distribution uniformity, the noise is denoised through the order and step size of the filter, and the speech signal at the position information of the speaker is enhanced.
[0036] In this embodiment, the filter order determines the complexity and processing ability of the filter. The higher the order, the more complex the signal changes that the filter can handle, but the computational amount will also increase accordingly. If the spectral range of the noise is narrow and mainly concentrated in the low-frequency or high-frequency region, a lower filter order can be selected to remove the noise at these specific frequencies. If the spectral range of the noise is wide or contains multiple noise components at different frequencies, a higher filter order may be needed to more effectively remove the noise.
[0037] In this embodiment, the step size parameter determines the speed and stability of the filter parameter update. When selecting, the change rate of the noise can be considered. For a rapidly changing noise environment, such as dynamic noise or burst noise, a larger step size can make the filter adapt to the change of the noise faster, but may lead to instability of the filtering result. For a slowly changing or relatively stable noise environment, a smaller step size can ensure the stability of the filtering result, but the speed at which the filter adapts to the change of the noise will be slower. On the spectrogram, the peak of the noise may change with the change of frequency. If the peak changes rapidly within a short time, it can be considered that the change rate of the noise is high. On the contrary, if the peak changes slowly or remains stable, the change rate of the noise may be low. By observing the distribution of energy on the spectrogram, we can have a certain understanding of the change rate of the noise. If the energy is evenly and stably distributed on the spectrogram, the change rate of the noise may be low. If the energy distribution is uneven and changes rapidly with time, the change rate of the noise may be high.
[0038] Step S102, extract multiple features of the speech signal, segment the speech signal to obtain multiple speech signal segments.
[0039] In this embodiment, the multiple features are multiple features on the spectral envelope diagram. By virtue of these features, the signal change situation is described, thereby segmenting the speech. The significance of segmentation is as follows. Improve recognition accuracy: Segmenting the speech signal can cut the continuous speech stream into smaller and relatively independent units, which usually correspond to words, phrases or sentences. Such segmentation helps the subsequent speech recognition system to more accurately recognize the content of each unit, because the speech features within each unit are more consistent and less affected by the context.
[0040] Reduce computational complexity: Performing recognition processing on the entire speech signal usually involves a huge amount of computation. By segmenting, the recognition task can be decomposed into multiple smaller tasks, and each task only processes one speech segment. This can significantly reduce the computational complexity and improve the recognition efficiency.
[0041] Facilitate subsequent processing: The segmented speech signal is easier to perform subsequent processing, such as the application of language models, speech synthesis, speech retrieval, etc. Each speech segment can be used as an independent processing unit, facilitating various operations and analyses.
[0042] Enhance robustness: In a noisy or complex speech environment, segmentation can help the recognition system better cope with noise and interference. By focusing on the feature changes within each speech segment, the system can more robustly recognize the speech content.
[0043] In some embodiments of the present application, multiple features of the speech signal are extracted, including Connect the highest amplitude points at different frequencies in the spectrogram of the speech signal to form a spectral envelope line, and perform smoothing processing to obtain a spectral envelope diagram; Uniformly set a number of marker points on the spectral envelope diagram, mark the peak points and valley points, and analyze the marker points, peak points and valley points on the spectral envelope diagram to calculate the feature parameters. The feature parameters include peak difference, valley difference, slope and entropy value. The peak difference is the amplitude difference between the marker point and its nearest peak point, and the valley difference is the amplitude difference between the marker point and its nearest valley point.
[0044] In this embodiment, the spectral envelope is the representation of a speech signal in the spectral domain, which reflects the spectral characteristics of the speech signal. The spectral envelope is a curve formed by connecting the highest amplitude points of different frequencies, and this curve is the spectral envelope line. The spectrum is a collection of many different frequencies, forming a very wide frequency range, and the amplitudes of different frequencies may be different. In speech signal processing, the spectral envelope reflects the spectral characteristics of the speech signal and is an important feature in speech signal processing. Observe the shape and change trend of the spectral envelope line and detect the significant change points of the spectral envelope. These significant change points usually correspond to the boundaries of different words or phrases in the speech signal because the spectral characteristics of different words or phrases are usually different.
[0045] It can be understood that setting multiple uniform marker points is to be able to comprehensively count the significant change points.
[0046] In this embodiment, the peak difference, valley difference, slope, and entropy value. The slope of the spectral envelope (i.e., the change rate of the envelope line at different frequency points) can reflect the rapid change of the energy distribution in the speech signal. When the slope changes significantly, it may mean that a certain word or phrase in the speech signal starts or ends, so it can be used as a reference for the segmentation boundary. The entropy value is an index to measure the complexity of the signal. In spectral envelope analysis, the change of the entropy value can reflect the change of the information amount in the speech signal. When the entropy value changes significantly, it may mean that an important part of the speech signal starts or ends, thus it can be used as a reference for the segmentation boundary.
[0047] The peak of the spectral envelope: The peak points in the spectral envelope usually correspond to the parts with the strongest energy in the speech signal, and these parts are often related to the vowels or certain consonants in the speech. The appearance and disappearance of the peak points can be used as important signs for determining the boundaries of speech segments.
[0048] The valley of the spectral envelope: Corresponding to the peak, the valley points in the spectral envelope represent the parts with lower energy, which may correspond to the silent segments or the weakened parts of consonants in the speech. The change of the valley can also provide clues for the determination of the segmentation boundary.
[0049] Peak difference, valley difference. On the spectral envelope curve, the peak and valley are important feature points, which represent the high and low points of the signal strength respectively. The distance or difference between the selected point and the peak and valley can reflect the positional relationship of this point relative to these feature points, thus providing additional information about the local characteristics of the signal. The distance or difference is a quantitative index that can be easily calculated and used for subsequent analysis. By incorporating these distances or differences into the calculation of the critical index, we can more comprehensively describe the characteristics of each point on the spectral envelope curve, thereby improving the accuracy of segmentation boundary recognition or the precision of speech feature analysis.
[0050] In some embodiments of the present application, the voice signal is segmented to obtain multiple voice signal segments, including calculating the fluctuation index of the marked points on the spectral envelope diagram based on the peak difference, valley difference, slope, and entropy value, and segmenting the voice signal through the fluctuation index to obtain multiple voice signal segments; ; wherein, is the fluctuation index of the th marked point, is the fluctuation conversion coefficient, , are respectively the peak difference and valley difference of the th marked point, represents the minimum value of the peak difference and valley difference of the th marked point, , are respectively the combined conversion coefficients of the slope and entropy value, is the slope of the th marked point, is a preset constant.
[0051] In this embodiment, represents the correction of the critical situation represented by the slope and entropy value to the peak difference or valley difference, so as to obtain an accurate critical index.
[0052] Step S103: Convert the multiple voice signal segments into multiple text content segments, analyze the correlation degree of the text content segments, and correlate the multiple text content segments to obtain semantic content.
[0053] In some embodiments of the present application, converting the multiple voice signal segments into multiple text content segments and analyzing the correlation degree of the text content segments includes inputting each voice signal segment into a preset acoustic model, outputting the text content segment corresponding to each voice signal segment, and removing redundant content from each text content segment, where the redundant content includes HTML tags and special characters; performing word segmentation on each text content segment and annotating the part-of-speech of each word after word segmentation, parsing the text content segment through a syntactic analyzer, and identifying the components in the text content segment and the relationships between the components; analyzing the part-of-speech coverage degree and the tightness of the component relationships of each text content segment, and defining the integrity and missing information of the text content segment in combination with the part-of-speech coverage degree and the tightness of the component relationships; when the integrity of the text content segment exceeds a preset integrity threshold, the text content segment is used as a complete text segment, otherwise, the text content segment is used as an incomplete text segment; For an incomplete text segment, compare the text content segments in the context of the incomplete text segment based on the missing information to check for a match, and splice the incomplete text segment based on the text content segments in the context to convert it into a complete text segment; Analyze the complete text segment and the incomplete text segment to define the degree of association.
[0054] In this embodiment, an acoustic model (such as a Hidden Markov Model HMM, a Deep Neural Network DNN, etc.) converts a speech signal into a corresponding language content segment, i.e., in text form. Removing noise: Noise refers to information that is irrelevant to the text theme or interferes with subsequent analysis, such as HTML tags, special characters, insignificant punctuation marks, etc. Methods for removing noise include regular expression matching, filtering out irrelevant words, using specialized cleaning tools, etc. Word segmentation is the process of splitting text into independent words and is a basic step in Chinese text processing. Word segmentation methods include dictionary-based word segmentation, statistics-based word segmentation (such as models like HMM, CRF, etc.), and word segmentation that combines dictionary and statistical methods. Part-of-speech tagging is the process of assigning grammatical attributes to each word, such as nouns, verbs, adjectives, etc. Part-of-speech tagging helps to understand the grammatical role of words in a sentence and provides a basis for subsequent syntactic analysis. Preprocessing can eliminate interference information in the text and improve the text quality; word segmentation and part-of-speech tagging provide a more refined text representation for subsequent analysis and help to deeply understand the text content.
[0055] In this embodiment, a syntactic analyzer is used to parse a sentence to identify the components of the sentence (such as the subject, predicate, object, etc.) and the relationships between the components (such as the subject-predicate relationship, the verb-object relationship, etc.). A syntactic analyzer is usually constructed based on theories such as Context-Free Grammar (CFG), Dependency Grammar (DG), or Combinatory Categorial Grammar (CCG). Define the integrity and missing information of the text content segment in combination with the part-of-speech coverage and the tightness of the component relationships. Part-of-speech coverage: Examine whether the text segment contains all the parts of speech required to form a complete syntactic structure. Analyze whether the syntactic relationships between the various components in the text segment are tight, coherent, and in the correct order. Combine the two to define the integrity of the text. Missing information refers to whether this text lacks components and which components are missing, etc. Then splice the text according to the context to make the text relatively complete.
[0056] In some embodiments of the present application, analyze the complete text segment to define the degree of association, including, Extract the key features in the complete text segment and the incomplete text segment, perform quantization and standardization processing on each key feature, construct the respective sentence feature sets for each complete text segment and each incomplete text segment, and calculate the degree of association; ; Among them, For the th text segment and the The correlation degree between the sentence feature set of the th text segment, denotes the th text segment and the adjustment factor, represents the adjustment factor obtained by mapping the feature with the greatest contribution in the intersection of the sentence feature sets, and respectively denote the th text segment and the th text segment's respective weights, and respectively denote the th
[0057] In this embodiment, multiple text contents may be similar or of the same theme, and they are associated to improve the understanding accuracy. By associating multiple text segments, the robot can more accurately understand the user's overall intention and specific requirements, avoiding misunderstandings or omissions. It enhances task coherence. The associated semantic content makes the robot's actions more coherent and orderly, improving the efficiency and fluency of task execution.
[0058] In this embodiment, syntactic analysis: First, perform syntactic analysis on each sentence to determine the structure, components, and their relationships of the sentence. This is the basis for understanding the sentence semantics. Content analysis: On the basis of syntactic analysis, further perform content analysis. This includes identifying synonyms, antonyms, hyponymy relationships, etc. Through content analysis, we can more accurately grasp the deep meaning of each sentence. According to the results of content analysis, calculate the association situation between these texts and extract the key features of each sentence. These features can be words, phrases, syntactic structures, or semantic relationships, etc. The selection of features should be based on their contribution degree to the sentence semantics.
[0059] In this embodiment, the correlation degree is obtained by improving the Jaccard similarity. An adjustment factor is determined by the feature with the greatest contribution in the intersection, represents the correction of the similarity by the integrity of the two texts (the higher the comprehensive integrity, the higher the credibility). The weights of the text segments are different for complete text segments and incomplete text segments.
[0060] In some embodiments of the present application, multiple text segments are associated to obtain semantic content, including Classify multiple text content segments based on the correlation between pairwise text content segments, extract the core semantic content of each category, and obtain the semantic content.
[0061] In this embodiment, according to the correlation calculation result, classify similar or related paragraphs. Summarize each category and extract the core semantic content of this category. Organize these core semantic contents in a logical order to form an overall semantic content framework.
[0062] For example, the issuer says: "Robot, please first walk to the window in the living room, then help me pull the curtain halfway, then come back to me, and finally tell me the weather today." In this example, there is a spatial continuity between "first walk to the window in the living room" and "then pull the curtain" (which can be associated through the correlation), "then come back to me" forms a sequential relationship in time with the previous actions, and "finally tell me the weather" is a supplement and conclusion to the previous series of actions.
[0063] Step S104, extract multiple action behaviors from the semantic content, integrate the action behaviors to generate a control instruction, and manipulate the humanoid robot to perform corresponding activities through the control instruction.
[0064] In this embodiment, extract specific action behavior instructions from the constructed semantic content. By integrating and optimizing these action behaviors, we generate the final control instruction. After receiving the instruction, the robot will perform corresponding activities such as walking, grasping an object, answering a question, etc., so as to achieve intelligent control based on hearing.
[0065] Compared with the prior art, the beneficial effects of the present invention are: 1. Analyze the noise situation of the voice signal and the position information of the issuer, perform denoising processing on the voice signal according to the noise situation and the position information of the issuer, enhance the voice signal specifically, and effectively remove the noise, improving the reliability of the voice signal data, providing a basis for subsequent voice analysis and recognition.
[0066] 2. Segment the voice signal to obtain multiple voice signal segments, analyze the correlation of the text content segments, and correlate multiple text content segments to obtain the semantic content. Segment the voice according to the characteristics of the voice signal, and perform text content association according to the correlation of the text content segments to obtain accurate semantics, thereby improving the robot's recognition ability for long voices. Improve the accuracy and effect of robot control, and ensure the stability of robot command execution.
[0067] Correspondingly, the present application also provides an auditory-based humanoid robot control system, as Figure 2 shown, including, A denoising module, which is used to collect the voice signal of the issuer, analyze the noise situation and the position information of the issuer of the voice signal, and perform denoising processing on the voice signal according to the noise situation and the position information of the issuer; A segmentation module, which is used to extract multiple features of the voice signal, segment the voice signal to obtain multiple voice signal segments; An association module, which is used to convert multiple voice signal segments into multiple text content segments, analyze the association degree of the text content segments, and associate the multiple text content segments to obtain semantic content; A control module, which is used to extract multiple action behaviors from the semantic content, integrate the action behaviors to generate a control instruction, and manipulate the humanoid robot to perform corresponding activities through the control instruction.
[0068] Through the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by hardware or by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various implementation scenarios of the present invention.
[0069] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0070] Those skilled in the art can understand that the modules in the system in the implementation scenario can be distributed in the system of the implementation scenario according to the description of the implementation scenario, or can be correspondingly changed and located in one or more systems different from this implementation scenario. The modules in the above implementation scenario can be combined into one module, or can be further split into multiple sub-modules.
[0071] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A humanoid robot control method based on hearing, characterized in that: include, Collect the voice signal of the sender, analyze the noise condition of the voice signal and the location information of the sender, and perform denoising on the voice signal according to the noise condition and the location information of the sender; Extract multiple features of the speech signal, segment the speech signal, and obtain multiple speech signal segments; Converting multiple speech signal segments into multiple text content segments, analyzing the relevance of the text content segments, and associating the multiple text content segments to obtain semantic content; Multiple action behaviors are extracted from the semantic content, and control instructions are generated by integrating the action behaviors. The humanoid robot is manipulated to perform corresponding activities through the control instructions.
2. The humanoid robot control method based on hearing according to claim 1, characterized in that: Analyze the noise level of the speech signal and send out the person's location information, including, Obtain the speech signal data containing noise in the past, extract the modeling features, determine the M value based on the AIC rule and the BIC rule, initialize the parameters of the Gaussian component, iteratively optimize the parameters of the Gaussian component, and train the GMM model; The GMM model is used to identify the noise of the speech signal, and the noise spectrum is obtained. The noise situation is confirmed on the noise spectrum, and the location information of the sender is determined by the sound source localization technology. The noise condition includes the spectrum range width, spectrum peak change rate and energy distribution uniformity of the noise.
3. The humanoid robot control method based on hearing according to claim 2, characterized in that: De-noising the speech signal based on the noise level and the sender's location information, including: The order of the filter is set based on the spectral range width of the noise, and the step size of the filter is set based on the spectral peak change rate and the uniformity of energy distribution. The noise is denoised by the order and step size of the filter, and the voice signal on the sender's location information is enhanced.
4. The method for controlling a humanoid robot based on hearing according to claim 1, characterized in that: Extract multiple features of speech signals, including, Connect the highest amplitude points of different frequencies in the spectrum of the speech signal to form a spectrum envelope, and perform smoothing to obtain a spectrum envelope diagram; A number of marking points are evenly set on the spectrum envelope diagram, and the peak points and valley points are marked. The marking points, peak points and valley points on the spectrum envelope diagram are analyzed to calculate the characteristic parameters. The characteristic parameters include peak difference, peak-valley difference, slope and entropy value. The peak difference is the difference in amplitude between the marking point and its nearest peak point, and the peak-valley difference is the difference in amplitude between the marking point and its nearest peak-valley point.
5. The method for controlling a humanoid robot based on hearing according to claim 4, characterized in that: The speech signal is segmented to obtain multiple speech signal segments. include, The fluctuation index of the marking point on the spectrum envelope diagram is calculated based on the peak difference, the peak-to-valley difference, the slope and the entropy value, and the speech signal is segmented by the fluctuation index to obtain multiple speech signal segments; ; in, For the The volatility indicator of the marked points, is the volatility conversion coefficient, , They are The peak difference and peak-to-valley difference of the marker points, Indicates The minimum value of the peak difference and the peak-to-valley difference of the marker points. , are the combined conversion coefficients of slope and entropy, For the The slope of the marked points, is a preset constant.
6. The method for controlling a humanoid robot based on hearing according to claim 1, characterized in that: Convert multiple speech signal segments into multiple text content segments, and analyze the relevance of the text content segments, including: Input each speech signal segment into a preset acoustic model, output the text content segment corresponding to each speech signal segment, and remove redundant content from each text content segment, including HTML tags and special characters; Each text content segment is segmented, and the part of speech of each word after segmentation is marked. The text content segment is parsed through a syntactic analyzer to identify the components in the text content segment and the relationship between the components; Analyze the part-of-speech coverage and component relationship closeness of each text content segment, and define the completeness and missing information of the text content segment based on the part-of-speech coverage and component relationship closeness; When the completeness of the text content segment exceeds a preset completeness threshold, the text content segment is regarded as a complete text segment; otherwise, the text content segment is regarded as an incomplete text segment; For an incomplete text segment, the text content segment of the context of the incomplete text segment is compared with the missing information to see whether it matches, and the incomplete text segment is spliced with the text content segment of the context to convert it into a complete text segment; Complete and incomplete text segments are analyzed to define relevance.
7. The method for controlling a humanoid robot based on hearing according to claim 6, characterized in that: Analyze the complete text segment to define relevance, including, Extract key features from complete text segments and incomplete text segments, quantify and standardize each key feature, construct a sentence feature set for each complete text segment and each incomplete text segment, and calculate the relevance; ; in, For the paragraphs and The correlation between the text segments, It is The sentence feature set of a text segment, It is The sentence feature set of a text segment, Indicates paragraphs and The feature that contributes most to the intersection of the sentence feature sets of the text segments, Adjustment factor, The adjustment factor obtained by mapping a feature with the largest contribution in the intersection of the sentence feature sets, , Respectively represent paragraphs and The weight of each text segment is , Respectively represent paragraphs and The completeness of each text segment, is a preset constant.
8. The method for controlling a humanoid robot based on hearing according to claim 1, characterized in that: And associate multiple text content segments to obtain semantic content, including, The multiple text content segments are classified according to the correlation between each two text content segments, and the core semantic content of each category is extracted to obtain the semantic content.
9. A humanoid robot control system based on hearing, characterized in that: include, A denoising module is used to collect the voice signal of the sender, analyze the noise condition of the voice signal and the location information of the sender, and denoise the voice signal according to the noise condition and the location information of the sender; A segmentation module is used to extract multiple features of the speech signal and segment the speech signal to obtain multiple speech signal segments; An association module, used to convert multiple speech signal segments into multiple text content segments, analyze the association of the text content segments, and associate the multiple text content segments to obtain semantic content; The control module is used to extract multiple action behaviors from the semantic content, integrate the action behaviors to generate control instructions, and control the humanoid robot to perform corresponding activities through the control instructions.
Citation Information
Cited By
Humanoid robot hearing device
CN122210708A