Voice conversation interaction method and system for industrial equipment

By enhancing voice signals through microphone arrays and incremental adaptive filtering algorithms, combined with information gain transfer learning and time delay estimation methods, the accuracy problem of voice interaction in industrial noise environments is solved, and efficient and accurate voice dialogue interaction is achieved.

CN120708615AActive Publication Date: 2025-09-26CHINA APPLIED TECH CO LTD

Patent Information

Application Number
CN202511179953.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-22
Publication Date
2025-09-26
Estimated Expiration
2045-08-22

AI Technical Summary

Technical Problem

In industrial noisy environments, existing speech enhancement technologies and speech recognition methods are unable to effectively process complex and changeable voice commands, resulting in insufficient accuracy and reliability of voice interaction, affecting user experience.

Method used

A microphone array and incremental adaptive filtering algorithm are used for speech enhancement, and the information gain transfer learning method is combined to identify the interaction intention. The time delay estimation method is used to identify the position, and a speaker output strategy is constructed to realize voice dialogue interaction.

Benefits of technology

It improves voice quality, increases the accuracy of interaction intent recognition, enables smooth and natural voice dialogue interaction, and enhances convenience and accuracy in industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708615A_ABST
    Figure CN120708615A_ABST
Patent Text Reader

Abstract

The invention provides a voice dialogue interaction method and system for industrial equipment, and relates to the technical field of man-machine interaction, and the method comprises the following steps: obtaining a voice signal of a user, and carrying out the voice enhancement processing of the voice signal of the user through an incremental adaptive filtering algorithm, and obtaining an enhanced voice; according to the enhanced voice, using an information gain transfer learning method to identify an interaction intention of the user, and generating a to-be-interacted voice based on the interaction intention of the user; based on the enhanced voice, utilizing a time delay estimation method to identify an interaction position of the user, and performing position prediction on the position of the user; according to the position prediction result, an industrial equipment horn output strategy is constructed; and outputting the to-be-interacted voice based on the industrial equipment loudspeaker output strategy so as to realize voice dialogue interaction between the user and the industrial equipment. According to the invention, convenience, high efficiency and accuracy of interaction between the user and the equipment in an industrial scene are greatly improved, and intelligent development of industrial production is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction technology, and in particular to a voice dialogue interaction method and system for industrial equipment. Background Art

[0002] In industrial production, with the continuous development of intelligent technologies, achieving efficient and convenient human-machine interaction has become a key step in improving production efficiency and quality. As a natural and efficient method of human-machine interaction, voice interaction has enormous potential for application in industrial equipment. It allows operators to communicate with equipment through voice commands, eliminating the need for manual operation and significantly improving operational convenience and flexibility.

[0003] When it comes to voice signal acquisition, industrial sites often experience high levels of ambient noise, such as the sounds of machinery and tools. This noise can severely interfere with the acquisition of user voice signals, resulting in reduced voice quality and making subsequent speech recognition and intent understanding difficult. While existing speech enhancement technologies can suppress noise to a certain extent, their effectiveness remains unsatisfactory in complex and variable industrial noise environments, making it difficult to meet the accuracy and reliability requirements for voice interaction in industrial equipment. Regarding interaction intent recognition, voice commands in the industrial field are highly specialized and semantically complex. Voice commands used by different industrial equipment and production processes vary significantly, and user expression habits vary widely, making accurate identification of user interaction intent a major challenge. Traditional speech recognition and intent understanding methods often struggle to adapt to the specificities of industrial scenarios and are unable to effectively handle these complex and variable voice commands. This results in low intent recognition accuracy, impacting the practicality and user experience of voice interaction.

[0004] Currently, no effective solutions have been proposed for the problems in related technologies. Summary of the Invention

[0005] In view of this, the present invention provides a voice dialogue interaction method and system for industrial equipment to solve the above-mentioned problems.

[0006] In order to solve the above problems, the specific technical solutions adopted by the present invention are as follows: According to one aspect of the present invention, a voice dialogue interaction method for industrial equipment is provided, comprising the following steps: S1. Based on the microphone array pre-configured on the industrial equipment, the user's voice signal is obtained and the user's voice signal is enhanced by voice enhancement processing using an incremental adaptive filtering algorithm to obtain enhanced voice. S2. Based on the enhanced speech, the information gain transfer learning method is used to identify the user's interaction intention and generate the speech to be interacted based on the user's interaction intention; S3. Based on enhanced speech, the user's interaction location is identified using a time delay estimation method, and the user's location is predicted. Based on the location prediction results, an industrial equipment speaker output strategy is constructed. S4. Based on the industrial equipment speaker output strategy, the voice to be interacted with is output to realize voice dialogue interaction between the user and the industrial equipment.

[0007] Preferably, the method of acquiring a user's voice signal based on a microphone array pre-configured on industrial equipment and performing voice enhancement processing on the user's voice signal through an incremental adaptive filtering algorithm to obtain enhanced voice includes the following steps: S11. Synchronously collecting spatially distributed multi-channel voice signals using a microphone array pre-configured on industrial equipment, and preprocessing the voice signals of each channel to obtain preprocessed voice signals; S12. constructing an initial error signal based on the preprocessed speech signal, and performing preliminary enhancement processing on the preprocessed speech signal using an incremental adaptive filtering algorithm to obtain a preliminary enhanced speech signal; S13. Using a fixed-time interference observer, perform transient noise estimation and compensation on the preliminary enhanced speech signal to obtain the final enhanced speech.

[0008] Preferably, the process of constructing an initial error signal based on the preprocessed speech signal and performing preliminary enhancement processing on the preprocessed speech signal using an incremental adaptive filtering algorithm to obtain a preliminary enhanced speech signal comprises the following steps: S121. For the pre-processed speech signals of each channel, randomly select a group of channels as reference channels, and construct an initial error signal based on the pre-processed speech signals of the reference channels; S122, initializing the parameters of the incremental adaptive filtering algorithm, including the filter order, step size factor, and initial filter coefficient; S123 , according to the initialized incremental adaptive filtering algorithm, using the initial error signal to iteratively update the filter coefficient, and filtering the speech signal after the reference channel preprocessing to obtain a preliminary enhanced speech signal.

[0009] Preferably, constructing the initial error signal according to the speech signal preprocessed by the reference channel comprises the following steps: The speech signal after reference channel preprocessing is divided into frames, and each frame of speech signal is processed by fast Fourier transform to obtain the spectrum of the speech signal; Based on the spectrum of the speech signal, the power spectrum of each frequency point is calculated. At each frequency point, a minimum value search is performed on each frequency point within a preset time sliding window to obtain the minimum power value of each frequency point within the time sliding window as the initial noise power spectrum estimate; Smoothing the initial estimated values ​​of noise power at each frequency point, and synthesizing the noise signal based on the smoothed noise power estimated values; The initial error signal is obtained by subtracting the synthesized noise signal from the reference channel speech signal.

[0010] Preferably, the method of identifying the user's interaction intention based on the enhanced speech and generating the speech to be interacted based on the user's interaction intention by using the information gain transfer learning method includes the following steps: S21. Convert the enhanced speech into a speech signal using speech recognition technology to obtain speech text; S22. Using the information gain method to extract semantic features from the obtained speech text, and using a forward learning algorithm based on acoustic semantic distance to process the extracted features to identify the user's interaction intention; S23. Based on the user's interaction intention, combined with the knowledge graph, speech synthesis technology is used to generate the speech to be interacted.

[0011] Preferably, the method of extracting semantic features from the obtained speech text using the information gain method and processing the extracted features using a forward learning algorithm based on acoustic semantic distance to identify the user's interaction intention includes the following steps: S221: performing text processing on the obtained speech text, and performing category statistics on the text-processed speech text based on preset industrial equipment category labels to determine the frequency of occurrence of each category; S222, determining the distribution of each word in different categories, and calculating the importance score of each word using an information gain algorithm based on the frequency of occurrence of each category, and extracting semantic features based on the importance score of each word; S223. Based on the semantic features of the words, the acoustic semantic distance between the words is calculated, and a semantic similarity matrix in the industrial scenario is constructed based on the acoustic semantic distance; S224. Utilize the transfer learning algorithm and combine it with the semantic similarity matrix to dynamically adjust the classification weights and output the user intent category.

[0012] Preferably, the step of calculating the acoustic semantic distance between words based on the semantic features of the words and constructing a semantic similarity matrix in an industrial scenario according to the acoustic semantic distance includes the following steps: S2331. For the semantic features of words, the feature decoupling technology is used to divide the semantic features into acoustic features and text features; S2332, extracting the acoustic feature vector and text semantic feature vector of each word based on the acoustic features and text features, and generating a semantic embedding vector through a fusion mechanism; S2333. Based on the fused semantic embedding vector, the cosine similarity is used to calculate the acoustic semantic distance between words, and a semantic similarity matrix is ​​constructed based on the calculated acoustic semantic distance.

[0013] Preferably, the method of utilizing a transfer learning algorithm, combining a semantic similarity matrix to dynamically adjust classification weights, and outputting user intent categories includes the following steps: S2241. Initialize the pre-trained intent classification model weights and use the intent classification model weights as the starting point for transfer learning. S2242. Constructing a graph Laplacian matrix according to the semantic similarity matrix, and constructing a graph Laplacian regularization term using the graph Laplacian matrix; S2243. Optimizing the weights of the intent classification model based on maximizing the likelihood function and combining the graph Laplace regularization term to obtain an optimized intent classification model. S2244. Obtain new voice text, use the optimized intent classification model to convert the new voice text into a feature vector, calculate the probability of each intent category, and select the category with the highest probability as the final user interaction intention.

[0014] Preferably, the method of identifying the user's interaction position based on enhanced speech using a time delay estimation method and predicting the user's position; and constructing an industrial equipment speaker output strategy based on the position prediction result includes the following steps: S31. Based on the enhanced speech, calculate the time delay estimate between the microphone pairs using the generalized cross-correlation-phase transformation method; S32. Obtain the coordinate position of the microphone array in space, and based on the estimated time delay between the microphone pairs, solve the estimated coordinates of the user's sound source by minimizing the sound source localization error function to identify the user's current interaction position; S33. Predict the user's next interaction location based on the historical interaction location and the current recognition result, obtain the target area, and construct a speaker output strategy according to the location of the target area.

[0015] According to another aspect of the present invention, a voice dialogue interaction system for industrial equipment is provided, the system comprising: The voice enhancement processing module is used to obtain the user's voice signal based on the microphone array pre-configured on the industrial equipment, and perform voice enhancement processing on the user's voice signal through an incremental adaptive filtering algorithm to obtain enhanced voice; The interaction intention recognition module is used to identify the user's interaction intention based on the enhanced speech and use the information gain transfer learning method, and generate the speech to be interacted based on the user's interaction intention; The strategy building module is used to identify the user's interaction location based on enhanced speech and use the time delay estimation method to predict the user's location; and build the industrial equipment speaker output strategy based on the location prediction results; The voice interaction module is used to output the voice to be interacted with based on the industrial equipment speaker output strategy, so as to realize the voice dialogue interaction between the user and the industrial equipment.

[0016] The beneficial effects of the present invention are: The present invention acquires and enhances user voice signals through a microphone array and an incremental adaptive filtering algorithm, which can effectively improve voice quality and reduce environmental noise interference; utilizes the information gain transfer learning method to identify interaction intentions and generate speech to be interacted with, which can more accurately understand user needs and improve the accuracy of interaction intention recognition; uses the time delay estimation method to identify and predict the interaction position, and then constructs a speaker output strategy to achieve targeted and accurate voice output; finally, the speech to be interacted is output according to the strategy, achieving smooth and natural voice dialogue interaction, greatly improving the convenience, efficiency and accuracy of user interaction with equipment in industrial scenarios, and promoting the intelligent development of industrial production. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1 is a flow chart of a voice dialogue interaction method for industrial equipment according to an embodiment of the present invention; Figure 2 This is a principle block diagram of a voice dialogue interaction system for industrial equipment according to an embodiment of the present invention.

[0018] In the picture: 1. Speech enhancement processing module; 2. Interaction intention recognition module; 3. Strategy construction module; 4. Speech interaction module. DETAILED DESCRIPTION

[0019] In order to enable those skilled in the art to better understand the technical solutions in this application, the following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0020] According to an embodiment of the present invention, a voice dialogue interaction method and system for industrial equipment are provided.

[0021] The present invention will now be further described with reference to the accompanying drawings and specific embodiments. Figure 1 As shown, according to one embodiment of the present invention, a voice dialogue interaction method for industrial equipment is provided, comprising the following steps: S1. Based on the microphone array pre-configured on the industrial equipment, the user's voice signal is obtained and the user's voice signal is enhanced by voice enhancement processing using an incremental adaptive filtering algorithm to obtain enhanced voice. As a preferred embodiment, the method of acquiring the user's voice signal based on a microphone array pre-configured on the industrial equipment and performing voice enhancement processing on the user's voice signal through an incremental adaptive filtering algorithm to obtain enhanced voice includes the following steps: S11. Synchronously collecting spatially distributed multi-channel voice signals using a microphone array pre-configured on industrial equipment, and preprocessing the voice signals of each channel to obtain preprocessed voice signals; It's important to note that preprocessing of the voice signal for each channel includes DC removal and pre-emphasis. DC removal eliminates DC offsets by calculating the signal's average value and subtracting it from each sampling point, allowing the signal to fluctuate around zero. Pre-emphasis uses a first-order high-pass filter to enhance the high-frequency component and flatten the signal spectrum, facilitating subsequent spectral analysis and feature extraction. Preprocessing improves signal quality and provides more suitable input for subsequent processing.

[0022] S12. constructing an initial error signal based on the preprocessed speech signal, and performing preliminary enhancement processing on the preprocessed speech signal using an incremental adaptive filtering algorithm to obtain a preliminary enhanced speech signal; As a preferred embodiment, the initial error signal is constructed based on the preprocessed speech signal, and the preprocessed speech signal is preliminarily enhanced using an incremental adaptive filtering algorithm to obtain a preliminary enhanced speech signal, including the following steps: S121. For the pre-processed speech signals of each channel, randomly select a group of channels as reference channels, and construct an initial error signal based on the pre-processed speech signals of the reference channels; Specifically, in a microphone array, the signal of any channel contains user voice and noise components; by randomly selecting a group of channels as reference channels, the signal of the channel can be used to estimate the noise and then construct an initial error signal.

[0023] As a preferred embodiment, the construction of the initial error signal based on the speech signal preprocessed by the reference channel includes the following steps: The speech signal after reference channel preprocessing is divided into frames, and each frame of speech signal is processed by fast Fourier transform to obtain the spectrum of the speech signal; It's important to note that a speech signal represented in the time domain cannot directly reflect its frequency components and energy distribution. Frequency domain analysis, on the other hand, can more intuitively demonstrate the signal's frequency characteristics, facilitating noise estimation and speech enhancement. The Fast Fourier Transform (FFT) is an efficient algorithm for calculating the Discrete Fourier Transform (DFT), rapidly converting time-domain signals into frequency-domain signals.

[0024] Specifically, the fast Fourier transform regards each frame of the time domain speech signal after framing as a discrete sequence, and converts it into a complex sequence in the frequency domain through the fast Fourier transform algorithm; the modulus of the complex sequence represents the amplitude of the signal at different frequencies, and the phase represents the phase information of the signal at different frequencies.

[0025] Based on the spectrum of the speech signal, the power spectrum of each frequency point is calculated. At each frequency point, a minimum value search is performed on each frequency point within a preset time sliding window to obtain the minimum power value of each frequency point within the time sliding window as the initial noise power spectrum estimate; It should be noted that for the frequency domain complex sequence after fast Fourier transform, its power spectrum can be obtained by calculating the square of the modulus of each frequency point. k frequency points, its power spectrum P ( k )for: P ( k )=∣ X ( k )∣ 2 ; in, X ( k ) is the first k The plural number of frequencies.

[0026] In addition, the searched minimum power value can be approximately regarded as the noise power at this frequency point, thereby obtaining an initial noise power spectrum estimation.

[0027] Smoothing the initial estimated values ​​of noise power at each frequency point, and synthesizing the noise signal based on the smoothed noise power estimated values; Specifically, the moving average method is used to smooth the initial estimated value of the noise power at each frequency point, and the smoothed noise power estimated value is obtained by averaging the noise power estimated value at the current frequency point with the estimated values ​​at several previous and subsequent frequency points.

[0028] When synthesizing noise signals, the frequency-domain noise power spectrum can be converted into a time-domain noise signal through an inverse fast Fourier transform (IFFT). This involves generating a frequency-domain noise complex sequence based on the noise power estimate (usually assuming the noise phase is random), and then performing an IFFT transform on the frequency-domain noise complex sequence to obtain a time-domain noise signal.

[0029] The initial error signal is obtained by subtracting the synthesized noise signal from the reference channel speech signal.

[0030] S122, initializing the parameters of the incremental adaptive filtering algorithm, including the filter order, step size factor, and initial filter coefficient; It's important to note that the filter order refers to the number of past input signal samples used in the adaptive filter; it determines the extent to which the filter utilizes historical information about the input signal. In incremental adaptive filtering algorithms, the higher the filter order, the wider the range of input signal history the filter can consider. The step size factor is a key parameter in incremental adaptive filtering algorithms, controlling the magnitude of the filter coefficient adjustments during each iterative update. The initial filter coefficients are the values ​​of the adaptive filter coefficients before the iterative update begins; they provide a starting point for the iterative update of the filter coefficients.

[0031] S123 , according to the initialized incremental adaptive filtering algorithm, using the initial error signal to iteratively update the filter coefficient, and filtering the speech signal after the reference channel preprocessing to obtain a preliminary enhanced speech signal.

[0032] Specifically, the core idea of ​​the incremental adaptive filtering algorithm is to continuously adjust the filter coefficients based on the input signal and the error signal, so that the filter output is as close as possible to the desired signal (in speech enhancement, the desired signal is a pure speech signal, but since it cannot be directly obtained, it is usually adjusted indirectly through the error signal). In each iteration, the algorithm updates the filter coefficients according to certain rules based on the current input signal, error signal, and step size factor. Specifically, the following steps are involved: The speech signal (input signal) preprocessed by the reference channel is convolved with the current filter coefficient to obtain the output signal of the filter.

[0033] The initial error signal (at the first iteration) or the error signal obtained from the previous iteration is compared with the filter output signal to obtain the error signal at the current moment. In speech enhancement, the error signal can be expressed as the difference between the desired signal and the filter output; Input signal and step size factor, update the filter coefficient according to the update formula of incremental adaptive filtering algorithm, the update formula is: ; in, m represents the step size factor, w i ( n ) indicates the time n When the adaptive filter i filter coefficients, w i ( n +1) means at time n +1, the adaptive filter i filter coefficients, e ( n ) indicates the time n The error signal.

[0034] In addition, after each iteration, the filter coefficients are updated, and the speech signal after reference channel preprocessing is filtered using the updated filter coefficients; the purpose of filtering is to remove noise components in the speech signal and obtain a signal that is closer to pure speech.

[0035] Specifically, in each iteration, the updated filter coefficients are convolved with the current input signal (the speech signal after reference channel preprocessing) to produce the filter output signal at that moment. As the iterations progress, the filter continuously adjusts its coefficients to better adapt to changes in the signal and noise, gradually improving the filtering effect. After multiple iterations, the filter output signal can be considered a preliminarily enhanced speech signal.

[0036] S13. Using a fixed-time interference observer, perform transient noise estimation and compensation on the preliminary enhanced speech signal to obtain the final enhanced speech.

[0037] It should be noted that a fixed-time interference observer is a device that can quickly and accurately estimate interference in a system (in speech signal processing scenarios, transient noise). Based on the system's input and output information and a preset model structure, it uses a specific algorithm to achieve interference estimation within a fixed time. Specifically, when using a fixed-time interference observer to estimate and compensate for transient noise in a preliminary enhanced speech signal, the following steps are included: Receive the preliminary enhanced speech signal, collect the signal in real time, and analyze the time domain and frequency domain characteristics of the signal; in the time domain, observe the amplitude change, waveform characteristics, etc. of the signal; in the frequency domain, convert the signal to the frequency domain through algorithms such as fast Fourier transform (FFT) to analyze the frequency component distribution of the signal.

[0038] Based on a preset system model and disturbance estimation algorithm, a fixed-time disturbance observer begins to estimate transient noise. For example, a fixed-time disturbance observer designed using sliding mode control theory defines an appropriate sliding surface and sliding mode control law, enabling the observer's state to converge to the actual disturbance value within a fixed time, thereby obtaining an estimated value of the transient noise. After obtaining the estimated value of transient noise, the weighted compensation method is used to supplement it. The weighted compensation method is to assign a weight coefficient to the noise estimate according to the accuracy of the noise estimation and the characteristics of the signal, and then use the weighted coefficient to perform weighted compensation.

[0039] S2. Based on the enhanced speech, the information gain transfer learning method is used to identify the user's interaction intention and generate the speech to be interacted based on the user's interaction intention; As a preferred embodiment, the method of identifying the user's interaction intention based on the enhanced speech and generating the speech to be interacted based on the user's interaction intention by using the information gain transfer learning method includes the following steps: S21. Convert the enhanced speech into a speech signal using speech recognition technology to obtain speech text; It's important to note that speech recognition technology converts the lexical content of human speech into a computer-readable text format; this process involves multiple key components, including acoustic models and language models. First, the acoustic model maps speech signal features (such as Mel-frequency cepstral coefficients) to phonemes or subword units. Trained on a large amount of speech data, it learns the probabilistic relationship between different speech features and corresponding pronunciations. Then, based on grammatical rules and a large text corpus, the language model combines and adjusts the phoneme or subword sequences output by the acoustic model to generate text that conforms to linguistic conventions.

[0040] S22. Using the information gain method to extract semantic features from the obtained speech text, and using a forward learning algorithm based on acoustic semantic distance to process the extracted features to identify the user's interaction intention; As a preferred embodiment, the method of extracting semantic features from the obtained speech text using the information gain method and processing the extracted features using a forward learning algorithm based on acoustic semantic distance to identify the user's interaction intention includes the following steps: S221: performing text processing on the obtained speech text, and performing category statistics on the text-processed speech text based on preset industrial equipment category labels to determine the frequency of occurrence of each category; It should be noted that text processing operations include: removing stop words, stemming, lemmatization, etc. Specifically, the speech text can be preprocessed through natural language processing tool libraries (such as NLTK, spaCy, etc.); first, the text is divided into words or phrases, then stop words are removed, and then stemming or lemmatization operations are performed to obtain the processed text.

[0041] In addition, the preset industrial equipment category labels are defined according to the actual needs and equipment types in industrial scenarios, such as machine tools, robots, conveyor belts, etc.

[0042] S222, determining the distribution of each word in different categories, and calculating the importance score of each word using an information gain algorithm based on the frequency of occurrence of each category, and extracting semantic features based on the importance score of each word; It should be noted that information gain is an indicator used to describe the importance of a feature to a classification system. It is based on the concept of information entropy and evaluates the importance of a feature by calculating the change in the system's information entropy when the feature exists and does not exist. The greater the information gain, the greater the contribution of the feature to the classification system.

[0043] For each word, count the number of times it appears in each category and calculate its frequency of occurrence in each category. Based on the frequency of occurrence in each category, calculate the information entropy of the probability distribution of all categories; then, for each word, calculate the conditional information entropy when the word exists and does not exist; finally, calculate the information gain value of each word as its importance score according to the information gain calculation formula; the information gain calculation formula is: information gain = system information entropy - conditional information entropy.

[0044] S223. Based on the semantic features of the words, the acoustic semantic distance between the words is calculated, and a semantic similarity matrix in the industrial scenario is constructed based on the acoustic semantic distance; As a preferred embodiment, the method of calculating the acoustic semantic distance between words based on the semantic features of the words and constructing a semantic similarity matrix in an industrial scenario according to the acoustic semantic distance includes the following steps: S2331. For the semantic features of words, the feature decoupling technology is used to divide the semantic features into acoustic features and text features; It should be noted that in scenarios where voice and text are combined, semantic features are a mixture of acoustic information from the user's voice (such as pitch, speaking speed, pronunciation, etc.) and semantic information from the text (such as the meaning of vocabulary, grammatical structure, etc.). Feature decoupling technology can separate features from these two different sources; specifically, by inputting the semantic features of vocabulary into a pre-trained feature decoupling model; the model will analyze and process the input features, and gradually separate the acoustic features and text features through internal network structure and parameter adjustment.

[0045] S2332, extracting the acoustic feature vector and text semantic feature vector of each word based on the acoustic features and text features, and generating a semantic embedding vector through a fusion mechanism; It should be noted that acoustic feature vectors are used to characterize the phonetic characteristics of words. Commonly used acoustic feature extraction methods include Mel-Frequency Cepstral Coefficients (MFCC) and Linear Predictive Coding (LPC). MFCC can effectively capture the spectral information in speech signals by simulating the human ear's perception of sound. Text semantic feature vectors are used to represent the semantic meaning of words. Common methods include word embedding technologies such as Word2Vec and GloVe. By learning the contextual relationships between words in a large text corpus, each word is mapped into a low-dimensional vector space, so that semantically similar words are closer in the vector space.

[0046] Specifically, after extracting the acoustic feature vector and the text semantic feature vector, the acoustic feature vector and the text semantic feature vector can be directly spliced ​​and fused to finally obtain a semantic embedding vector.

[0047] S2333. Based on the fused semantic embedding vector, the cosine similarity is used to calculate the acoustic semantic distance between words, and a semantic similarity matrix is ​​constructed based on the calculated acoustic semantic distance.

[0048] It should be noted that cosine similarity measures the degree of directional similarity between two vectors. Its value range is [−1, 1]. Values ​​closer to 1 indicate more directional similarity between the two vectors; values ​​closer to -1 indicate more directional opposites between the two vectors; and a value of 0 indicates that the two vectors are orthogonal, meaning unrelated. When calculating the acoustic-semantic distance between words, cosine similarity can effectively measure the similarity between the semantic embedding vectors of two words.

[0049] S224. Utilize the transfer learning algorithm and combine it with the semantic similarity matrix to dynamically adjust the classification weights and output the user intent category.

[0050] As a preferred embodiment, the method of using a transfer learning algorithm, combining a semantic similarity matrix to dynamically adjust classification weights, and outputting user intent categories includes the following steps: S2241. Initialize the pre-trained intent classification model weights and use the intent classification model weights as the starting point for transfer learning. It should be noted that in industrial scenarios, intent classification models can be selected from deep learning models pre-trained on industrial speech and text data, such as models based on recurrent neural networks (RNN) and its variants (long short-term memory networks (LSTMs) and gated recurrent units (GRUs)), or models based on Transformer architectures (such as BERT and its variants).

[0051] The core idea of ​​transfer learning is to transfer knowledge learned from a pre-trained model to a new task. Using the pre-trained model weights as a starting point, this existing knowledge can be leveraged in new industrial intent classification tasks, reducing training data requirements and improving the model's convergence speed and generalization capabilities.

[0052] S2242. Constructing a graph Laplacian matrix according to the semantic similarity matrix, and constructing a graph Laplacian regularization term using the graph Laplacian matrix; It should be noted that when constructing the graph Laplacian matrix, the words are treated as nodes in the graph, and the edges between the nodes are determined based on the semantic similarity matrix. If the semantic similarity between two words is above a certain threshold, an edge is established between them, and the weight of the edge can be set to their semantic similarity value.

[0053] Among them, the graph Laplacian matrix L The adjacency matrix of the graph can be A Sum degree matrix D The calculation formula is Among them, the degree matrix D is a diagonal matrix whose diagonal elements D ii Representation node i degree, that is, the degree of the node i The sum of the weights of the connected edges. The graph Laplacian matrix can capture the local structure and relationships between nodes in the graph.

[0054] In addition, the purpose of constructing the graph Laplace regularization term is to introduce semantic similarity information during the model training process, so that the model can take into account the semantic relationship between words when classifying. By adding the regularization term to the loss function, the parameter update of the model can be constrained so that the model output corresponding to similar words is also similar; the graph Laplace regularization term R ( i ) can be expressed as R ( i )= i T Lθ .in, i T is a parameter i The transpose of .

[0055] S2243. Optimizing the weights of the intent classification model based on maximizing the likelihood function and combining the graph Laplace regularization term to obtain an optimized intent classification model. It should be noted that by maximizing the likelihood function, we can find the parameters that maximize the probability that the model will output the correct label for a given data set. This is done by taking the logarithm of the likelihood function to obtain the log-likelihood function, which is then maximized using an optimization algorithm (such as stochastic gradient descent (SGD) or Adam).

[0056] Among them, the graph Laplace regularization term is added to the log-likelihood function to construct a new loss function J ( i ) is: ; in, l is a hyperparameter that controls the strength of the regularization term. l The larger it is, the stronger the constraint effect of the regularization term on the update of model parameters; l The smaller it is, X is the input data, Y is the output label, and the model is more likely to fit the training data; then an optimization algorithm (such as the Adam algorithm) is used to minimize the loss function J ( i ). In each iteration, the loss function is calculated with respect to the model parameters i The gradient of is calculated, and the parameters are updated according to the gradient. Through continuous iteration, the loss function is gradually reduced, and the model parameters gradually converge to the optimal value, thus obtaining the optimized intent classification model.

[0057] S2244. Obtain new voice text, use the optimized intent classification model to convert the new voice text into a feature vector, calculate the probability of each intent category, and select the category with the highest probability as the final user interaction intention.

[0058] It should be noted that the feature vector is input into the classification layer of the optimized intent classification model. The classification layer is a fully connected layer followed by a softmax function. The softmax function converts the output of the classification layer into a probability distribution for each intent category.

[0059] S23. Based on the user's interaction intention, combined with the knowledge graph, speech synthesis technology is used to generate the speech to be interacted.

[0060] It should be noted that before using speech synthesis technology to generate the interactive voice based on the user's interaction intention and in combination with the knowledge graph, it is necessary to build a knowledge graph containing multiple contents such as device information, operating procedures, fault knowledge, product descriptions, etc.

[0061] Based on the user's intent, information is retrieved from the knowledge graph. Using a graph database query language or a graph-based retrieval method, nodes and edges relevant to the user's intent are found. For example, if the user's intent is to inquire about the operating procedures for a particular device, the device node is found in the knowledge graph. Specific operational information is then retrieved along the edges related to the operating procedures. Finally, the retrieved information is integrated into a complete and coherent response.

[0062] S3. Based on enhanced speech, the user's interaction location is identified using a time delay estimation method, and the user's location is predicted. Based on the location prediction results, an industrial equipment speaker output strategy is constructed. As a preferred embodiment, the method of identifying the user's interaction location based on enhanced speech and using a time delay estimation method and predicting the user's location; and constructing an industrial equipment speaker output strategy based on the location prediction result includes the following steps: S31. Based on the enhanced speech, calculate the time delay estimate between the microphone pairs using the generalized cross-correlation-phase transformation method; It should be noted that the cross-correlation function is used to measure the similarity between two signals at different time offsets. x 1( n )and x 2( n ), their cross-correlation function R 12 ( t ) is defined as: ; in, t is the time offset, when R 12 ( t ) obtains the maximum value, the corresponding t The value is the estimated delay between the two signals.

[0063] In addition, in order to improve the accuracy of delay estimation and anti-interference ability, generalized cross-correlation is introduced. It highlights the useful information in the signal and suppresses the influence of noise and reverberation by weighting the cross-correlation function. Phase transformation is a commonly used weighting method for generalized cross-correlation. Its weighting function W ( oh )for: ; in,X 1( oh )and X 2( oh ) are x 1( n )and x 2( n ) Fourier transform, phase transformation weighting can highlight the phase information of the signal and reduce the influence of amplitude information on delay estimation. It can achieve better delay estimation effect in a strong reverberation environment. By calculating the generalized cross-correlation function after phase transformation weighting, we can find the maximum value corresponding to t The delay between the microphone pair can be estimated by using the value of

[0064] S32. Obtain the coordinate position of the microphone array in space based on the estimated time delay between the microphone pairs; solve the estimated coordinates of the user's sound source by minimizing the sound source localization error function, and identify the user's current interaction position; It should be noted that after the coordinate position of each microphone in the microphone array and the estimated value of the time delay between microphone pairs are known, the speed of sound propagation in air (usually 340 m / s ) to establish the geometric relationship between the delay estimate and the coordinates.

[0065] From a geometric perspective, using microphones at known coordinates as reference points, the path lengths of sound traveling from the user's sound source to different microphones are different. The delay estimate actually reflects the difference between these path lengths. For example, if there are two microphones and sound arrives at one first and then the other, the delay estimate between them corresponds to the difference in the sound's path lengths between the two microphones.

[0066] When there are multiple microphones in the microphone array, multiple such microphone pairs will be formed. Each pair can reflect a geometric constraint relationship between the user's sound source position and these microphones based on the time delay estimation value and the sound propagation speed. By combining the geometric constraint information provided by all microphone pairs, a complete conditional framework can be constructed to determine the estimated position coordinates of the user's sound source in space. Then, by minimizing the sound source localization error function, the estimated position coordinates of the user's sound source can be solved, and finally the user's current interaction position can be identified.

[0067] S33. Predict the user's next interaction location based on the historical interaction location and the current recognition result, obtain the target area, and construct a speaker output strategy according to the location of the target area.

[0068] It should be noted that when predicting the user's next interaction location and constructing a speaker output strategy based on historical interaction locations and current recognition results, the user's interaction location information over a period of time will be continuously recorded. Then, combined with the patterns shown by the historical interaction locations and the current location, a specific analysis method will be used to infer the range of interaction locations where the user is most likely to go next. This range is the target area. After determining the target area, it is necessary to construct a speaker output strategy based on its location. First, based on the relative position relationship between the target area and each speaker, the speaker that can most effectively cover the area will be selected. If the target area is close to a specific speaker, this speaker can be used for voice output first, and the volume of the speaker output can be adjusted according to the size of the target area and the user's possible activities to ensure that the user can clearly hear the voice prompts in the target area without interfering with the normal work of other areas due to excessive volume.

[0069] S4. Based on the industrial equipment speaker output strategy, the voice to be interacted with is output to realize voice dialogue interaction between the user and the industrial equipment.

[0070] It should be noted that, based on the previously established speaker output strategy, the interactive voice is optimally broadcast through the industrial equipment's speakers, enabling voice interaction between the user and the device, ensuring that the voice information reaches the user accurately and clearly. This process not only ensures the effective transmission of information.

[0071] like Figure 2 As shown, according to one embodiment of the present invention, a voice dialogue interaction system for industrial equipment is provided, the system comprising: The voice enhancement processing module 1 is used to obtain the user's voice signal based on the microphone array pre-configured on the industrial equipment, and perform voice enhancement processing on the user's voice signal through an incremental adaptive filtering algorithm to obtain enhanced voice; Interaction intention recognition module 2 is used to identify the user's interaction intention based on the enhanced speech and use the information gain transfer learning method, and generate the speech to be interacted based on the user's interaction intention; Strategy building module 3 is used to identify the user's interaction location based on enhanced speech and use the time delay estimation method to predict the user's location; and to build an industrial equipment speaker output strategy based on the location prediction results; The voice interaction module 4 is used to output the voice to be interacted with based on the industrial equipment speaker output strategy, so as to realize the voice dialogue interaction between the user and the industrial equipment.

[0072] To sum up, with the help of the above technical solutions of the present invention, the present invention obtains and enhances user voice signals through microphone arrays and incremental adaptive filtering algorithms, which can effectively improve voice quality and reduce environmental noise interference; uses the information gain transfer learning method to identify interaction intentions and generate voices to be interacted with, which can more accurately understand user needs and improve the accuracy of interaction intention recognition; uses the time delay estimation method to identify the interaction position and make predictions, and then constructs a speaker output strategy to achieve targeted and accurate voice output; finally, the voice to be interacted is output according to the strategy, achieving smooth and natural voice dialogue interaction, greatly improving the convenience, efficiency and accuracy of user interaction with equipment in industrial scenarios, and contributing to the intelligent development of industrial production.

[0073] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, optical storage, etc.) containing computer-usable program code.

[0074] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A voice dialogue interaction method for industrial equipment, characterized in that: The following steps are involved: S1. Based on the microphone array pre-configured on the industrial equipment, the user's voice signal is obtained and the user's voice signal is enhanced by voice enhancement processing using an incremental adaptive filtering algorithm to obtain enhanced voice. S2. Based on the enhanced speech, the information gain transfer learning method is used to identify the user's interaction intention and generate the speech to be interacted based on the user's interaction intention; S3. Based on the enhanced speech, the user's interaction location is identified using the time delay estimation method, and the user's location is predicted; And based on the position prediction results, build the industrial equipment speaker output strategy; S4. Based on the industrial equipment speaker output strategy, the voice to be interacted with is output to realize voice dialogue interaction between the user and the industrial equipment.

2. The voice dialogue interaction method for industrial equipment according to claim 1, characterized in that: The method of acquiring a user's voice signal based on a microphone array pre-configured on the industrial equipment and performing voice enhancement processing on the user's voice signal through an incremental adaptive filtering algorithm to obtain enhanced voice includes the following steps: S11. Synchronously collecting spatially distributed multi-channel voice signals using a microphone array pre-configured on industrial equipment, and preprocessing the voice signals of each channel to obtain preprocessed voice signals; S12. constructing an initial error signal based on the preprocessed speech signal, and performing preliminary enhancement processing on the preprocessed speech signal using an incremental adaptive filtering algorithm to obtain a preliminary enhanced speech signal; S13. Using a fixed-time interference observer, perform transient noise estimation and compensation on the preliminary enhanced speech signal to obtain the final enhanced speech.

3. The voice dialogue interaction method for industrial equipment according to claim 2, characterized in that: The method of constructing an initial error signal based on the preprocessed speech signal and performing preliminary enhancement processing on the preprocessed speech signal using an incremental adaptive filtering algorithm to obtain a preliminary enhanced speech signal includes the following steps: S121. For the pre-processed speech signals of each channel, randomly select a group of channels as reference channels, and construct an initial error signal based on the pre-processed speech signals of the reference channels; S122, initializing the parameters of the incremental adaptive filtering algorithm, including the filter order, step size factor, and initial filter coefficient; S123 , according to the initialized incremental adaptive filtering algorithm, using the initial error signal to iteratively update the filter coefficient, and filtering the speech signal after the reference channel preprocessing to obtain a preliminary enhanced speech signal.

4. The voice dialogue interaction method for industrial equipment according to claim 3, characterized in that: The method of constructing an initial error signal based on the speech signal preprocessed by the reference channel comprises the following steps: The speech signal after reference channel preprocessing is divided into frames, and each frame of speech signal is processed by fast Fourier transform to obtain the spectrum of the speech signal; Based on the spectrum of the speech signal, the power spectrum of each frequency point is calculated. At each frequency point, a minimum value search is performed on each frequency point within a preset time sliding window to obtain the minimum power value of each frequency point within the time sliding window as the initial noise power spectrum estimate; Smoothing the initial estimated values ​​of noise power at each frequency point, and synthesizing the noise signal based on the smoothed noise power estimated values; The initial error signal is obtained by subtracting the synthesized noise signal from the reference channel speech signal.

5. The voice dialogue interaction method for industrial equipment according to claim 2, characterized in that: The method of identifying the user's interaction intention based on the enhanced speech and generating the speech to be interacted based on the user's interaction intention by using the information gain transfer learning method includes the following steps: S21. Convert the enhanced speech into a speech signal using speech recognition technology to obtain speech text; S22. Using the information gain method to extract semantic features from the obtained speech text, and using a forward learning algorithm based on acoustic semantic distance to process the extracted features to identify the user's interaction intention; S23. Based on the user's interaction intention, combined with the knowledge graph, speech synthesis technology is used to generate the speech to be interacted.

6. The voice dialogue interaction method for industrial equipment according to claim 5, characterized in that: The method of extracting semantic features from the obtained speech text using the information gain method and processing the extracted features using a forward learning algorithm based on acoustic semantic distance to identify the user's interaction intention includes the following steps: S221: performing text processing on the obtained speech text, and performing category statistics on the text-processed speech text based on preset industrial equipment category labels to determine the frequency of occurrence of each category; S222, determining the distribution of each word in different categories, and calculating the importance score of each word using an information gain algorithm based on the frequency of occurrence of each category, and extracting semantic features based on the importance score of each word; S223. Based on the semantic features of the words, the acoustic semantic distance between the words is calculated, and a semantic similarity matrix in the industrial scenario is constructed based on the acoustic semantic distance; S224. Utilize the transfer learning algorithm and combine it with the semantic similarity matrix to dynamically adjust the classification weights and output the user intent category.

7. The voice dialogue interaction method for industrial equipment according to claim 6, characterized in that: The method of calculating the acoustic semantic distance between words based on the semantic features of the words and constructing a semantic similarity matrix in an industrial scenario based on the acoustic semantic distance includes the following steps: S2331. For the semantic features of words, the feature decoupling technology is used to divide the semantic features into acoustic features and text features; S2332, extracting the acoustic feature vector and text semantic feature vector of each word based on the acoustic features and text features, and generating a semantic embedding vector through a fusion mechanism; S2333. Based on the fused semantic embedding vector, the cosine similarity is used to calculate the acoustic semantic distance between words, and a semantic similarity matrix is ​​constructed based on the calculated acoustic semantic distance.

8. The voice dialogue interaction method for industrial equipment according to claim 6, characterized in that: The method of utilizing the transfer learning algorithm, combining the semantic similarity matrix to dynamically adjust the classification weight, and outputting the user intent category includes the following steps: S2241. Initialize the pre-trained intent classification model weights and use the intent classification model weights as the starting point for transfer learning. S2242. Constructing a graph Laplacian matrix according to the semantic similarity matrix, and constructing a graph Laplacian regularization term using the graph Laplacian matrix; S2243. Optimizing the weights of the intent classification model based on maximizing the likelihood function and combining the graph Laplace regularization term to obtain an optimized intent classification model. S2244. Obtain new voice text, use the optimized intent classification model to convert the new voice text into a feature vector, calculate the probability of each intent category, and select the category with the highest probability as the final user interaction intention.

9. The voice dialogue interaction method for industrial equipment according to claim 1, characterized in that: Based on the enhanced speech, the user's interaction position is identified by using a time delay estimation method, and the user's position is predicted; Based on the position prediction results, the construction of industrial equipment speaker output strategy includes the following steps: S31. Based on the enhanced speech, calculate the time delay estimate between the microphone pairs using the generalized cross-correlation-phase transformation method; S32. Obtain the coordinate position of the microphone array in space, and based on the estimated time delay between the microphone pairs, solve the estimated coordinates of the user's sound source by minimizing the sound source localization error function to identify the user's current interaction position; S33. Predict the user's next interaction location based on the historical interaction location and the current recognition result, obtain the target area, and construct a speaker output strategy according to the location of the target area.

10. A voice dialogue interaction system for industrial equipment, used to implement the voice dialogue interaction method for industrial equipment according to any one of claims 1 to 9, characterized in that: The system includes: The voice enhancement processing module is used to obtain the user's voice signal based on the microphone array pre-configured on the industrial equipment, and perform voice enhancement processing on the user's voice signal through an incremental adaptive filtering algorithm to obtain enhanced voice; The interaction intention recognition module is used to identify the user's interaction intention based on the enhanced speech and use the information gain transfer learning method, and generate the speech to be interacted based on the user's interaction intention; The strategy building module is used to identify the user's interaction location based on enhanced speech and use the time delay estimation method to predict the user's location; and build the industrial equipment speaker output strategy based on the location prediction results; The voice interaction module is used to output the voice to be interacted with based on the industrial equipment speaker output strategy, so as to realize the voice dialogue interaction between the user and the industrial equipment.

Citation Information

Patent Citations

  • Voice interaction method and system

    CN108564943A

  • Text keyword extraction method and device, apparatus, and storage medium

    CN109635273A

  • Voice interaction method and device based on virtual robot image and vehicle-mounted equipment intelligent control system

    CN111124123A

  • Semantic vector representation method and device, computer equipment and storage medium

    CN116127981A

  • Voice interaction service system based on artificial intelligence

    CN116913277A

Cited By

  • Voice recognition software interaction implementation method based on workshop production operation scene

    CN120998183A

  • Industrial equipment control method and system based on voice interaction

    CN121415782A