Mobile phone voice interaction optimization method and system combined with environment perception

Through the processing method of combining ambient light information with microphone array and personalized sound feature library, the problem of speech recognition errors in noisy environments is solved, and the accuracy and user experience of speech recognition are improved.

CN120455584AInactive Publication Date: 2025-08-08深圳市图高智能有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510562151.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In noisy environments, background noise can easily interfere with microphone collection, resulting in speech recognition errors. The prior art is difficult to effectively suppress noise interference, resulting in misrecognition or false wake-up of voice assistants.

Method used

Ambient sound is collected through the microphone array, and the target user's personalized sound feature library configuration assist in judging features, generate a microphone signal suppression feature matrix, locate the main direction and secondary direction, perform noise suppression and signal enhancement, and adjust voice interaction parameters in combination with real-time ambient light information to realize the wake-up and interaction processing of the voice interaction module.

Benefits of technology

It significantly improves the accuracy of voice recognition, reduces misrecognition, optimizes the human-computer interaction experience, and ensures that the voice assistant can accurately recognize user voice in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455584A_ABST
    Figure CN120455584A_ABST
Patent Text Reader

Abstract

The invention provides a mobile phone voice interaction optimization method and system combined with environment perception, and relates to the technical field of voice recognition, and the method comprises the steps: carrying out the collection of environment sound through a microphone array of a target mobile phone; configuring auxiliary judgment features, and performing comparison and signal suppression analysis on the environment sound data set; positioning a main direction and an auxiliary direction based on the microphone signal suppression feature matrix; performing cooperative noise processing in the primary and secondary directions by using the primary direction voice data and the secondary direction voice data; and collecting real-time ambient light information and an ambient sound data set to execute voice interaction parameter constraint, and performing awakening of the voice interaction module or interaction processing after awakening. The technical problem of voice recognition errors caused by the fact that background noise in a noisy environment easily interferes microphone acquisition is solved, and voice recognition accuracy and response speed in a complex and changeable environment are improved through noise suppression and voice enhancement, so that man-machine interaction experience is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition technology, and in particular to a method and system for optimizing mobile phone speech interaction combined with environmental perception. Background Art

[0002] Mobile voice interaction refers to intelligent interaction with mobile phones through voice recognition, natural language processing, and speech synthesis, allowing users to control their phones through voice commands. Currently, voice assistants in smartphones can perform a variety of functions, typically relying on built-in microphones and speech recognition engines, which achieve a certain degree of accurate recognition and response to voice commands. However, with the continuous development of voice interaction technology, the limitations of traditional voice assistants have gradually become apparent. Especially in complex environments, the accuracy of voice recognition is affected by many factors, such as background noise and interference from multiple voices, which can cause voice signal distortion or inability to correctly recognize the target voice, resulting in speech recognition errors. Furthermore, when multiple people are talking or the environment changes, voice assistants can easily misidentify commands from others or even accidentally wake up the device, causing unnecessary operations.

[0003] In summary, the prior art has a technical problem in that background noise in a noisy environment easily interferes with microphone acquisition, resulting in speech recognition errors. Summary of the Invention

[0004] The purpose of this application is to provide a mobile phone voice interaction optimization method and system combined with environmental perception, so as to solve the technical problem in the prior art that background noise in a noisy environment easily interferes with microphone collection, resulting in voice recognition errors.

[0005] In view of the above problems, the present application provides a method and system for optimizing mobile phone voice interaction combined with environmental perception.

[0006] In the first aspect, the present application provides a method for optimizing mobile phone voice interaction in combination with environmental perception, which is implemented by a mobile phone voice interaction optimization system in combination with environmental perception, wherein the method for optimizing mobile phone voice interaction in combination with environmental perception includes: collecting environmental sounds through the microphone array of the target mobile phone to generate an environmental sound data set; configuring auxiliary judgment features with the user personality sound feature library of the target user pre-stored in the voice interaction module, comparing and signal suppression analysis of the environmental sound data set, and generating a microphone signal suppression feature matrix; after locating the main direction and the secondary direction based on the microphone signal suppression feature matrix and performing noise suppression control, collecting the main direction voice data and the secondary direction voice data; performing collaborative noise processing of the main and secondary directions with the main direction voice data and the secondary direction voice data to generate enhanced voice data; collecting the real-time ambient light information of the target user, and performing voice interaction parameter constraints in combination with the environmental sound data set, and then performing wake-up or post-wake-up interaction processing on the enhanced voice data of the target user based on the user personality sound feature library.

[0007] Optionally, the voiceprint features of the target user are constructed using sound feature samples in the user's personalized sound feature library, and the auxiliary judgment features are configured; the auxiliary judgment features are used to perform sound comparison in different directions on the ambient sound data of each microphone in the ambient sound data set to determine the main direction sound signal and the secondary direction sound signal; signal enhancement configuration is performed on the main direction sound signal, and noise suppression configuration is performed on the secondary direction sound signal to generate the microphone signal suppression feature matrix.

[0008] Optionally, performing signal enhancement configuration on the main direction sound signal includes maximizing a gain in the main direction.

[0009] Optionally, the auxiliary judgment feature is used as a signal separation standard to separate the secondary direction sound signal from the noise signal and determine the power spectrum of the noise signal; the power spectrum of the secondary direction sound signal is determined; the filter gain is calculated based on the power spectrum of the noise signal and the power spectrum of the secondary direction sound signal, and the secondary direction sound signal is configured for noise suppression using the filter gain.

[0010] Optionally, the auxiliary judgment feature is used as the main direction sound feature, and the main direction sound feature intensity and noise introduction intensity of the ambient sound data of each microphone are analyzed to generate the main sound intensity index and noise intensity index of each direction; based on the main sound intensity index and noise intensity index of each direction, the direction with the maximum main sound intensity and the minimum noise intensity is located, and the main direction sound signal and the secondary direction sound signal are separated from the ambient sound data of each microphone.

[0011] Optionally, a spectral model is constructed using the main direction speech data, and the speech spectrum in the secondary direction speech data is repaired to generate a secondary direction speech repair result; the noise spectrum in the secondary direction speech data is used to estimate and restore the noise of the main direction speech data to generate a main direction speech repair result; and the enhanced speech data is generated using the main direction speech repair result and the secondary direction speech repair result.

[0012] Optionally, the light characteristics and noise characteristics of multiple interactive environment modalities, as well as the corresponding preset interaction parameters, are determined; the real-time ambient light information and the ambient sound data set are identified using the light characteristics and noise characteristics of the multiple interactive environment modalities to determine the target interactive environment modality; and the voice interaction parameter constraints are completed using the preset interaction parameters corresponding to the target interactive environment modality.

[0013] Optionally, an interactive semantic recognition network is trained based on the user personality voice feature library; semantic recognition is performed on the target user enhanced voice data using the interactive semantic recognition network, and the voice interaction module is awakened or interactively processed after awakening using the semantic recognition result.

[0014] Optionally, if there are multiple rounds of interaction scenarios, the recognition is corrected based on the context of the multiple rounds of interaction scenarios.

[0015] In a second aspect, the present application also provides a mobile phone voice interaction optimization system combined with environmental perception, which is used to execute the mobile phone voice interaction optimization method combined with environmental perception as described in the first aspect, wherein the mobile phone voice interaction optimization system combined with environmental perception includes: an ambient sound acquisition module, which is used to collect ambient sound through the microphone array of the target mobile phone to generate an ambient sound data set; a signal suppression analysis module, which is used to configure auxiliary judgment features with the user personality voice feature library of the target user pre-stored in the voice interaction module, compare and perform signal suppression analysis on the ambient sound data set, and generate a microphone signal suppression feature matrix; a noise suppression control module, which is used to locate the main direction and the secondary direction based on the microphone signal suppression feature matrix and perform noise suppression control, and then collect the main direction voice data and the secondary direction voice data; a collaborative noise processing module, which is used to perform collaborative noise processing of the main and secondary directions with the main direction voice data and the secondary direction voice data to generate enhanced voice data; an interaction processing module, which is used to collect the real-time ambient light information of the target user, and perform voice interaction parameter constraints in combination with the ambient sound data set, and then perform wake-up or post-wake-up interaction processing on the enhanced voice data of the target user based on the user personality voice feature library.

[0016] One or more technical solutions provided in this application have at least the following beneficial effects:

[0017] Ambient sound is collected through the microphone array of the target mobile phone to generate an ambient sound data set; auxiliary judgment features are configured with the user personality sound feature library of the target user pre-stored in the voice interaction module, and the ambient sound data set is compared and signal suppression analysis is performed to generate a microphone signal suppression feature matrix; after locating the main direction and the secondary direction based on the microphone signal suppression feature matrix and performing noise suppression control, the main direction voice data and the secondary direction voice data are collected; collaborative noise processing of the main and secondary directions is performed with the main direction voice data and the secondary direction voice data to generate enhanced voice data; real-time ambient light information of the target user is collected, and voice interaction parameter constraints are performed in combination with the ambient sound data set, and then the voice interaction module is awakened or interactively processed after awakening on the enhanced voice data of the target user based on the user personality sound feature library. That is to say, the ambient sound is collected through the microphone array, and the auxiliary judgment features are configured according to the target user's personalized sound feature library to perform signal suppression analysis, locate the main direction and secondary direction and perform noise suppression, automatically enhance the main direction signal, suppress noise in other directions, improve voice clarity, and adjust the parameters of the voice interaction module according to the real-time ambient light information and ambient sound data set to realize wake-up or post-wake-up interaction processing, so that the mobile phone can automatically adapt to different environments, significantly improve the accuracy of voice recognition, reduce misrecognition, and thus optimize the human-computer interaction experience.

[0018] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, which can be implemented in accordance with the contents of the description, and to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are specifically listed below. It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in this application or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and a person of ordinary skill in the art can obtain other drawings based on the provided drawings without creative work.

[0020] Figure 1 This is a flow chart of the method for optimizing mobile phone voice interaction combined with environmental perception in this application;

[0021] Figure 2 This is a structural diagram of the mobile phone voice interaction optimization system combined with environmental perception in this application.

[0022] Explanation of the reference numerals: ambient sound collection module 11 , signal suppression analysis module 12 , noise suppression control module 13 , collaborative noise processing module 14 , interactive processing module 15 . DETAILED DESCRIPTION

[0023] This application solves the technical problem in the prior art of speech recognition errors caused by background noise in noisy environments easily interfering with microphone acquisition by providing a mobile phone voice interaction optimization method and system that combines environmental perception. Ambient sound is collected through a microphone array, and signal suppression analysis is performed based on auxiliary judgment features configured according to the target user's personalized voice feature library. The main direction and secondary direction are located and noise suppression is performed. The main direction signal is automatically enhanced, noise in other directions is suppressed, and speech clarity is improved. The parameters of the voice interaction module are adjusted according to real-time ambient light information and ambient sound data sets to achieve wake-up or post-wake-up interaction processing, allowing the mobile phone to automatically adapt to different environments, significantly improving the accuracy of speech recognition, reducing misrecognition, and thus optimizing the human-computer interaction experience.

[0024] Below, the technical solutions in this application will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited to the example embodiments described herein. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. It should also be noted that, for the convenience of description, only the parts related to this application, rather than all of them, are shown in the accompanying drawings.

[0025] For example, see the attached Figure 1 The present application provides a method for optimizing mobile phone voice interaction in combination with environment perception, wherein the method for optimizing mobile phone voice interaction in combination with environment perception is executed by a mobile phone voice interaction optimization system in combination with environment perception, and the method for optimizing mobile phone voice interaction in combination with environment perception specifically comprises the following steps:

[0026] S100: Collect ambient sound through the microphone array of the target mobile phone to generate an ambient sound dataset.

[0027] Specifically, the target mobile phone is usually equipped with multiple microphones to form a microphone array for collecting sound. The microphone array can capture sounds from different directions. For example, some high-end smartphones usually have at least 3 microphones, located at the bottom, top and back of the device, allowing the phone to build a more accurate environmental sound map by receiving audio signals from different directions. Through the microphone array, the mobile phone can collect sounds from different directions at the same time and generate an environmental sound dataset containing different sound sources, including background noise, human voices, traffic sounds and other audio signals.

[0028] S200: configuring auxiliary judgment features with a user personality voice feature library of a target user pre-stored in a voice interaction module, performing comparison and signal suppression analysis on the environmental sound data set, and generating a microphone signal suppression feature matrix.

[0029] Furthermore, the present application S200 includes:

[0030] The voiceprint features of the target user are constructed using the sound feature samples in the user's personalized sound feature library, and the auxiliary judgment features are configured; the auxiliary judgment features are used to perform sound comparison in different directions on the ambient sound data of each microphone in the ambient sound data set to determine the main direction sound signal and the secondary direction sound signal; signal enhancement configuration is performed on the main direction sound signal, and noise suppression configuration is performed on the secondary direction sound signal to generate the microphone signal suppression feature matrix.

[0031] Performing signal enhancement configuration on the main direction sound signal includes maximizing a gain in the main direction.

[0032] Using the auxiliary judgment feature as a signal separation standard, the secondary direction sound signal is separated from the noise signal to determine the power spectrum of the noise signal; the power spectrum of the secondary direction sound signal is determined; the filter gain is calculated based on the power spectrum of the noise signal and the power spectrum of the secondary direction sound signal, and the secondary direction sound signal is configured for noise suppression using the filter gain.

[0033] Specifically, within the voice interaction module, the target user's personalized voice feature library will store the user's unique voice attributes (such as voiceprint features, intonation, pronunciation, etc.), which is established by collecting voice samples of the target user in advance. For example, when a user first sets up the voice assistant, through a series of voice sampling (such as saying a few specific words or sentences), the spectral characteristics and pitch changes of the user's voice are extracted to form the user's personalized voice feature library. The more the user interacts with the phone, the richer the sound feature samples in the user's personalized voice feature library.

[0034] Based on the sound feature samples in the user's personalized sound feature library, a series of voiceprint features are extracted to construct the target user's voiceprint features. Voiceprint features are characteristics of the user's voice, including but not limited to the audio spectrum, pitch, timbre, pronunciation characteristics, etc., which can uniquely identify a person's voice. Voiceprint recognition technology is used to analyze the user's voice and extract their unique voiceprint features, such as Mel-Frequency Cepstral Coefficients (MFCC), Linear Prediction Cepstral Coefficients (LPCC), and spectrogram analysis. For example, MFCC is used to extract audio signal features by mimicking the human ear's response to sound frequency. By calculating the short-time Fourier transform of the sound, the spectral features of each frame are obtained. Then, a Mel-frequency transform is performed to ultimately obtain coefficients that can represent the user's voice characteristics, thereby obtaining the target user's specific voice features, including pitch, timbre, spectrum, etc., which together constitute the voiceprint feature.

[0035] The auxiliary judgment feature is used as the main direction sound feature, and the main direction sound feature intensity and noise introduction intensity of the ambient sound data of each microphone are analyzed, and the main direction sound signal and the secondary direction sound signal of each microphone are separated accordingly. The main direction sound signal is enhanced, that is, the directional gain is adjusted according to the target speaking direction. Gain is the degree of signal amplification, which is usually used to adjust the intensity of the audio signal and optimize the volume and clarity of the target signal according to the source direction of the signal. According to the direction of the target user's speech, the gain is adjusted to maximize the improvement of the sound signal from that direction. Through the formula G(θ)=max(G min ,G0*cos 2 (θ)) is calculated, where G(θ) = is the gain value of the target direction θ; G min It is a set minimum gain value to prevent the signal from being too weak when the gain is too low; G0 is the basic gain value, which is used to adjust the gain strength. θ is the angle between the target direction and the microphone array direction, cos(θ) 2 Indicates the angle relative to the target direction. The closer to the target direction, the larger the value, and the greater the gain. The gain is greatest for sound signals aligned with the target direction (θ), and decreases as the angle deviates.

[0036] The gain value of each direction is calculated according to the formula, and the signal received by each microphone is enhanced by the gain value. The signal from the target direction will be amplified, while the signal from the non-target direction will be weakened accordingly, enhancing the sound signal in the main direction and effectively reducing the noise interference from other directions. For example, suppose the target user is speaking in a noisy environment and the speaking direction is east, that is, 180°. Assuming G0 = 10dB, the minimum gain G min = 2dB (make sure the gain does not fall below this value to avoid the signal being too weak), then the gain is 10dB.

[0037] Using the auxiliary judgment features as signal separation criteria, the secondary direction sound signal is separated from the noise signal to determine the noise component in the secondary direction. In other words, the background noise in the secondary direction is separated and the power spectrum of the secondary signal is calculated. The power spectrum of the noise signal describes the energy distribution of the noise at different frequencies. The power spectrum is usually calculated using the Fourier transform method. Assuming that the noise signal is a signal in the time domain, the fast Fourier transform (FFT) is applied to convert it to the frequency domain, and then the power spectrum of the signal is calculated. The power spectrum is a frequency domain characteristic that describes the distribution of the signal at different frequencies and reflects the energy distribution of the signal at different frequency components.

[0038] Similarly, the power spectrum of the secondary direction sound signal is calculated, which reflects the distribution of the signal in the frequency domain and describes the intensity of different frequency components. When the power spectrum of the secondary direction signal and the noise signal is calculated, according to the filtering formula Calculate the filter gain to determine the attenuation of the noise signal in the secondary direction at each frequency. Where G(f) is the filter gain at frequency f, S xx (f) is the power spectrum of the secondary direction sound signal, S nn (f) is the power spectrum of the noise signal. Based on the calculated filter gain, noise suppression processing is performed on the secondary direction sound signal to reduce the noise components in these signals, thereby improving the signal-to-noise ratio of the speech signal.

[0039] After completing signal enhancement and noise suppression, a microphone signal suppression feature matrix is generated, recording the signal enhancement and noise suppression effects of each microphone in different directions. The microphone signal suppression feature matrix is calculated using a signal enhancement and noise suppression algorithm. It can optimize signal processing in a multi-microphone environment, ensuring that the voice signal in the primary direction is enhanced while the noise signal in the secondary direction is effectively suppressed. It dynamically identifies the user's speaking direction, automatically enhances the primary direction signal, suppresses noise in other directions, improves speech clarity, and reduces interference. By enhancing the primary direction signal and suppressing noise in the secondary direction, the clarity of the speech signal is significantly improved, and the interference of background noise on speech recognition is reduced, especially in noisy environments.

[0040] Furthermore, the present application further comprises the following steps:

[0041] Using the auxiliary judgment feature as the main direction sound feature, the main direction sound feature intensity and noise introduction intensity of the ambient sound data of each microphone are analyzed to generate the main sound intensity index and noise intensity index of each direction; based on the main sound intensity index and noise intensity index of each direction, the direction with the maximum main sound intensity and the minimum noise intensity is located, and the main direction sound signal and the secondary direction sound signal are separated from the ambient sound data of each microphone.

[0042] Specifically, auxiliary judgment features are configured according to the voiceprint features of the target user. The auxiliary judgment features are used as the main direction sound features, and the main direction sound feature intensity analysis is performed on the ambient sound data obtained from each microphone to identify the main direction sound features transmitted in different directions, that is, the sound features of the target user in each direction, and the main sound intensity index in each direction is obtained, which reflects the clarity and intensity of the target user's sound signal. Based on the main direction sound features, the ambient sound data obtained from each microphone is subjected to noise introduction intensity analysis, that is, the noise intensity introduced by the ambient noise (such as wind, traffic, other people's voices, etc.). The degree of noise interference is determined by analyzing the frequency components, sound patterns, etc., and a noise intensity index is generated. The noise intensity index refers to the intensity of the noise part in the ambient sound, that is, the intensity of the voice of the non-target user. It is quantified by analyzing the noise part (non-target user voice part) in the sound signal, and is usually manifested as an interference component in the sound signal.

[0043] Based on the primary sound intensity and noise intensity indicators for each direction, the direction with the highest primary sound intensity and lowest noise intensity is identified as the primary direction. The ambient sound data from each microphone is then separated to extract the primary direction sound signal and the secondary direction sound signal. The primary direction sound signal refers to the voice data collected from the direction where the target user's voice signal is the strongest and clearest among all directions. The secondary direction sound signal is the sound signal collected from non-primary directions, where the user's voice intensity is the weakest and the noise intensity is the highest. For example, in a certain direction, the user's speaking voice intensity is the highest and the noise intensity is the lowest. Therefore, using beamforming technology, the sound signal in this direction is separated as the primary direction sound signal, while the sound signals in other directions are considered secondary direction sound signals. By accurately separating the sound signals from the primary and secondary directions, the noise source is precisely located and the noise signal in the secondary direction is effectively suppressed, resulting in a clearer final voice signal.

[0044] S300: After locating the main direction and the secondary direction based on the microphone signal suppression feature matrix and performing noise suppression control, voice data in the main direction and voice data in the secondary direction are collected.

[0045] Specifically, according to the signal enhancement and noise suppression algorithm, the microphone signal suppression feature matrix is obtained, and the generated microphone signal suppression feature matrix is used to locate the main direction and the secondary direction, and the noise suppression control is started. In the main direction, the signal enhancement algorithm (such as gain adjustment) is used to enhance the voice signal of the target user; in the secondary direction, the noise suppression algorithm (such as gain filtering) is used to reduce the noise signal from the secondary direction. After noise suppression and signal enhancement, voice data is collected from the main direction and the secondary direction respectively. The voice data in the main direction will be enhanced to ensure that the user's voice signal has high clarity and accuracy; the voice data in the secondary direction will be processed with noise suppression to reduce the interference of irrelevant noise. By locating the main direction and the secondary direction, the voice signal of the target user is enhanced, and the noise is suppressed, the clarity and recognition accuracy of the voice signal are improved.

[0046] S400: Performing primary and secondary direction collaborative noise processing using the primary direction voice data and the secondary direction voice data to generate enhanced voice data.

[0047] Furthermore, the present application S400 includes:

[0048] A spectrum model is constructed using the main direction speech data, and the speech spectrum in the secondary direction speech data is repaired to generate a secondary direction speech repair result; noise estimation and restoration are performed on the main direction speech data using the noise spectrum in the secondary direction speech data to generate a main direction speech repair result; and the enhanced speech data is generated using the main direction speech repair result and the secondary direction speech repair result.

[0049] Specifically, frequency domain transformation (such as Fourier transform) is performed on the main direction voice data to extract the spectral characteristics of the signal and construct a spectral model. The spectral model describes the energy distribution of the signal at different frequencies. Based on the spectral model, the various components in the voice signal can be analyzed and repaired. According to the spectral model, the voice spectrum in the secondary direction voice data is repaired. The voice signal in the secondary direction loses clarity due to the interference of environmental noise. Spectral restoration technology is used to restore the voice components that are masked or distorted by noise. The secondary direction voice data usually contains strong environmental noise, so it is necessary to analyze the spectrum of this secondary direction data. By restoring the secondary direction voice signal through the spectral model, the noise in the secondary direction voice data is removed to a certain extent, and the voice components are retained to supplement the main direction audio.

[0050] Based on the noise spectrum in the secondary voice data—that is, the noise's representation in the frequency domain (energy representation in the frequency distribution)—noise estimation and restoration are performed on the primary voice data. This involves modeling and separating the noise component of the mixed signal based on known noise characteristics (such as the noise spectrum), suppressing minor noise in the primary voice data while retaining the clean user voice. After noise estimation and restoration, the primary voice signal is enhanced in clarity and can be used by speech recognition or interaction modules.

[0051] Enhanced speech data is generated by combining the primary and secondary restoration results. The independence of each track is preserved, with a primary direction and multiple secondary directions marked. In other words, the multiple restoration results for the primary and secondary directions are retained separately and not directly merged into a single track. Instead, they are used as multiple input channels for processing by the subsequent semantic recognition module to form a multi-channel enhanced speech collection.

[0052] Through spectral restoration and noise estimation, the system effectively removes noise, ensuring the clarity of voice data collected from multiple directions. This improves the accuracy and robustness of the voice interaction system in complex environments, enhancing the user experience. Specifically, it suppresses noise from the original voice in the primary direction to restore clean speech, while enhancing the user's voice in the secondary direction with weaker signal strength to remove background noise.

[0053] S500: Collect the real-time ambient light information of the target user, and execute voice interaction parameter constraints in combination with the ambient sound data set, and then perform wake-up or post-wake-up interaction processing on the target user's enhanced voice data based on the user's personalized voice feature library.

[0054] Furthermore, the present application S500 includes:

[0055] Determine the light characteristics and noise characteristics of multiple interactive environment modalities, as well as the corresponding preset interaction parameters; use the light characteristics and noise characteristics of the multiple interactive environment modalities to identify the real-time ambient light information and the ambient sound dataset to determine the target interactive environment modality; and complete the voice interaction parameter constraints using the preset interaction parameters corresponding to the target interactive environment modality.

[0056] Specifically, the target phone's built-in sensors collect ambient light information, including data such as light intensity, color temperature, and light and shadow distribution. For example, the data collected will differ significantly depending on whether the user is in sunlight outdoors, softly lit indoors, or dimly lit. Based on application requirements, the environment is divided into various interaction modes. For example, a noisy indoor environment has high background noise and bright light; a quiet indoor environment has low background noise and soft light; strong outdoor light is high brightness and relatively noisy; and a nighttime environment has low light and is relatively quiet.

[0057] Determine the light and noise characteristics of multiple interaction environment modalities and configure corresponding preset interaction parameters, such as wake-up word sensitivity, voice volume, and speech rate. For example, the preset parameters for outdoor bright light mode are low wake-up word sensitivity (to prevent false wake-ups), high voice volume (due to high ambient noise), and slow speech rate; the preset parameters for quiet indoor mode are high wake-up word sensitivity, moderate voice volume, and moderate speech rate.

[0058] Based on the collected real-time ambient light information and ambient sound data sets, the light characteristics and noise characteristics are analyzed to determine the matching target interaction environment modality among multiple interaction environment modalities. According to the preset interaction parameters corresponding to the target interaction environment modality, the voice interaction parameter constraints are completed to constrain and adjust the voice interaction process. For example, lower or increase the wake-up word detection sensitivity, adjust the output volume, speech speed, etc. By utilizing the real-time collected light and sound data, the interaction environment modality is automatically identified, ensuring that the mobile phone automatically adapts to different environments (noisy, quiet, outdoor, indoor, etc.), thereby improving the accuracy of voice recognition.

[0059] Furthermore, the present application further comprises the following steps:

[0060] An interactive semantic recognition network is trained based on the user's personalized voice feature library; semantic recognition is performed on the target user's enhanced voice data using the interactive semantic recognition network, and the voice interaction module is awakened or interactively processed after awakening using the semantic recognition result.

[0061] Specifically, based on the user's personalized voice feature library, we extract user-specific voice features (voiceprint, voice spectrum, intonation, etc.) and divide them into training and validation sets. We then design a deep neural network model and integrate personalized feature channels into the network, enabling the model to learn the target user's language habits in a customized manner. Supervised learning is performed on the training set, allowing the network to capture the user's personalized semantic features while converting speech to text, thereby improving the recognition accuracy and robustness of the target user's speech.

[0062] The enhanced voice data obtained after multi-microphone signal processing and noise suppression is input into the trained interactive semantic recognition network. The interactive semantic recognition network performs semantic recognition on the enhanced voice data of the target user, and performs wake-up of the voice interaction module or interactive processing after wake-up. For example, when "Hi, Google" or "Hey, Google Gemini" is recognized, the voice module is woken up. After wake-up, the corresponding operations are performed based on the recognition results and the predetermined strategy, such as querying the weather, sending messages, opening applications, etc. The predetermined strategy is a series of pre-set strategies for performing corresponding operations based on keywords. For example, a user says "Hi, Google, turn on the flashlight" to the device in an outdoor environment. The enhanced voice data generated after multi-microphone enhancement and noise suppression is input into the personalized interactive semantic recognition network. The network recognizes that the semantic result is to turn on the flashlight, activates the voice interaction module and the corresponding operation (such as starting the flashlight function).

[0063] The semantic recognition network trained using the user's personalized voice feature library can accurately identify the target user's accent, speaking speed, intonation and interaction habits, reducing recognition errors.

[0064] Furthermore, the present application further comprises the following steps:

[0065] If there are multiple rounds of interaction scenarios, the recognition is corrected based on the context of the multiple rounds of interaction scenarios.

[0066] Specifically, in multi-round interaction scenarios—that is, scenarios where a user and recipient engage in a continuous voice conversation, where each round of conversation relies not only on the current input but also on contextual information from previous rounds—the system leverages conversation history information or contextual semantic associations to correct or supplement the current speech recognition result. This includes determining the plausibility of the currently recognized text based on previously recognized conversation content and automatically correcting any unreasonable or ambiguous content. In multi-round interaction scenarios, it's necessary to understand the user's likely intent based on the context, compare it with the current recognition result, and then correct the current recognition result. For example, if "Open App" is recognized but contextual information is lacking, the user's intent can be further confirmed by asking, "Which app do you want to open?" For example, if the context indicates the user is discussing the weather, and the current recognition result contains text about a city, the result will automatically be adjusted to reflect the weather information for that city. This multi-round contextual correction eliminates ambiguities and errors in single-round recognition, improving overall recognition accuracy.

[0067] In summary, the mobile phone voice interaction optimization method combined with environmental perception provided by this application has the following advantages:

[0068] Beneficial effects:

[0069] Ambient sound is collected through the microphone array of the target mobile phone to generate an ambient sound data set; auxiliary judgment features are configured with the user personality sound feature library of the target user pre-stored in the voice interaction module, and the ambient sound data set is compared and signal suppression analysis is performed to generate a microphone signal suppression feature matrix; after locating the main direction and the secondary direction based on the microphone signal suppression feature matrix and performing noise suppression control, the main direction voice data and the secondary direction voice data are collected; collaborative noise processing of the main and secondary directions is performed with the main direction voice data and the secondary direction voice data to generate enhanced voice data; real-time ambient light information of the target user is collected, and voice interaction parameter constraints are performed in combination with the ambient sound data set, and then the voice interaction module is awakened or interactively processed after awakening on the enhanced voice data of the target user based on the user personality sound feature library. That is to say, the ambient sound is collected through the microphone array, and the auxiliary judgment features are configured according to the target user's personalized sound feature library to perform signal suppression analysis, locate the main direction and secondary direction and perform noise suppression, automatically enhance the main direction signal, suppress noise in other directions, improve voice clarity, and adjust the parameters of the voice interaction module according to the real-time ambient light information and ambient sound data set to realize wake-up or post-wake-up interaction processing, so that the mobile phone can automatically adapt to different environments, significantly improve the accuracy of voice recognition, reduce misrecognition, and thus optimize the human-computer interaction experience.

[0070] Example 2: Based on the same inventive concept as the mobile phone voice interaction optimization method combined with environment perception in the aforementioned Example 1, this application also provides a mobile phone voice interaction optimization system combined with environment perception, please refer to the attached Figure 2 The mobile phone voice interaction optimization system combined with environmental perception includes:

[0071] The ambient sound collection module 11 is used to collect ambient sound through the microphone array of the target mobile phone and generate an ambient sound data set; the signal suppression analysis module 12 is used to configure auxiliary judgment features with the user personality sound feature library of the target user pre-stored in the voice interaction module, compare and perform signal suppression analysis on the ambient sound data set, and generate a microphone signal suppression feature matrix; the noise suppression control module 13 is used to locate the main direction and the secondary direction based on the microphone signal suppression feature matrix and perform noise suppression control, and then collect the main direction voice data and the secondary direction voice data; the collaborative noise processing module 14 is used to perform collaborative noise processing of the main and secondary directions with the main direction voice data and the secondary direction voice data to generate enhanced voice data; the interaction processing module 15 is used to collect the real-time ambient light information of the target user, and execute voice interaction parameter constraints in combination with the ambient sound data set, and then perform wake-up of the voice interaction module or interaction processing after wake-up on the enhanced voice data of the target user based on the user personality sound feature library.

[0072] Furthermore, the signal suppression analysis module 12 in the mobile phone voice interaction optimization system combined with environment perception is also used for:

[0073] The voiceprint features of the target user are constructed using the sound feature samples in the user's personalized sound feature library, and the auxiliary judgment features are configured; the auxiliary judgment features are used to perform sound comparison in different directions on the ambient sound data of each microphone in the ambient sound data set to determine the main direction sound signal and the secondary direction sound signal; signal enhancement configuration is performed on the main direction sound signal, and noise suppression configuration is performed on the secondary direction sound signal to generate the microphone signal suppression feature matrix.

[0074] Furthermore, the signal suppression analysis module 12 in the mobile phone voice interaction optimization system combined with environment perception is also used for:

[0075] Performing signal enhancement configuration on the main direction sound signal includes maximizing a gain in the main direction.

[0076] Furthermore, the signal suppression analysis module 12 in the mobile phone voice interaction optimization system combined with environment perception is also used for:

[0077] Using the auxiliary judgment feature as a signal separation standard, the secondary direction sound signal is separated from the noise signal to determine the power spectrum of the noise signal; the power spectrum of the secondary direction sound signal is determined; the filter gain is calculated based on the power spectrum of the noise signal and the power spectrum of the secondary direction sound signal, and the secondary direction sound signal is configured for noise suppression using the filter gain.

[0078] Furthermore, the signal suppression analysis module 12 in the mobile phone voice interaction optimization system combined with environment perception is also used for:

[0079] Using the auxiliary judgment feature as the main direction sound feature, the main direction sound feature intensity and noise introduction intensity of the ambient sound data of each microphone are analyzed to generate the main sound intensity index and noise intensity index of each direction; based on the main sound intensity index and noise intensity index of each direction, the direction with the maximum main sound intensity and the minimum noise intensity is located, and the main direction sound signal and the secondary direction sound signal are separated from the ambient sound data of each microphone.

[0080] Furthermore, the collaborative noise processing module 14 in the mobile phone voice interaction optimization system combined with environment perception is also used for:

[0081] A spectrum model is constructed using the main direction speech data, and the speech spectrum in the secondary direction speech data is repaired to generate a secondary direction speech repair result; noise estimation and restoration are performed on the main direction speech data using the noise spectrum in the secondary direction speech data to generate a main direction speech repair result; and the enhanced speech data is generated using the main direction speech repair result and the secondary direction speech repair result.

[0082] Furthermore, the interaction processing module 15 in the mobile phone voice interaction optimization system combined with environment perception is also used for:

[0083] Determine the light characteristics and noise characteristics of multiple interactive environment modalities, as well as the corresponding preset interaction parameters; use the light characteristics and noise characteristics of the multiple interactive environment modalities to identify the real-time ambient light information and the ambient sound dataset to determine the target interactive environment modality; and complete the voice interaction parameter constraints using the preset interaction parameters corresponding to the target interactive environment modality.

[0084] Furthermore, the interaction processing module 15 in the mobile phone voice interaction optimization system combined with environment perception is also used for:

[0085] An interactive semantic recognition network is trained based on the user's personalized voice feature library; semantic recognition is performed on the target user's enhanced voice data using the interactive semantic recognition network, and the voice interaction module is awakened or interactively processed after awakening using the semantic recognition result.

[0086] Furthermore, the interaction processing module 15 in the mobile phone voice interaction optimization system combined with environment perception is also used for:

[0087] If there are multiple rounds of interaction scenarios, the recognition is corrected based on the context of the multiple rounds of interaction scenarios.

[0088] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. Figure 1 The mobile phone voice interaction optimization method combined with environmental perception and the specific examples in Example 1 are also applicable to the mobile phone voice interaction optimization system combined with environmental perception in this embodiment. Through the above detailed description of the mobile phone voice interaction optimization method combined with environmental perception, those skilled in the art can clearly understand the mobile phone voice interaction optimization system combined with environmental perception in this embodiment, so for the sake of brevity of the specification, it will not be described in detail here.

[0089] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

[0090] Obviously, for those skilled in the art, several improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the scope of protection of the present application.

Claims

1. A mobile phone voice interaction optimization method combined with environmental perception is characterized by: include: The ambient sound is collected through the microphone array of the target mobile phone to generate an ambient sound dataset; The auxiliary judgment feature is configured using the user personality voice feature library of the target user pre-stored in the voice interaction module, and the ambient sound data set is compared and the signal suppression analysis is performed to generate a microphone signal suppression feature matrix; After locating the main direction and the secondary direction based on the microphone signal suppression feature matrix and performing noise suppression control, collecting the main direction voice data and the secondary direction voice data; Performing primary and secondary direction collaborative noise processing on the primary direction voice data and the secondary direction voice data to generate enhanced voice data; The real-time ambient light information of the target user is collected, and voice interaction parameter constraints are implemented in combination with the ambient sound data set. Then, based on the user's personalized voice feature library, the target user's enhanced voice data is used to wake up the voice interaction module or perform interaction processing after wake-up.

2. The method for optimizing mobile phone voice interaction combined with environmental perception according to claim 1, characterized in that: The auxiliary judgment feature is configured using the user personality voice feature library of the target user pre-stored in the voice interaction module, and the ambient sound data set is compared and signal suppression analysis is performed to generate a microphone signal suppression feature matrix, including: Constructing the target user's voiceprint features with the voice feature samples in the user's individual voice feature library, and configuring the auxiliary judgment features; Performing sound comparison in different directions on the ambient sound data of each microphone in the ambient sound dataset using the auxiliary judgment feature to determine a main direction sound signal and a secondary direction sound signal; Signal enhancement configuration is performed on the main direction sound signal, and noise suppression configuration is performed on the secondary direction sound signal to generate the microphone signal suppression feature matrix.

3. The method for optimizing mobile phone voice interaction combined with environmental perception according to claim 2, characterized in that: Performing signal enhancement configuration on the main direction sound signal includes maximizing a gain in the main direction.

4. The method for optimizing mobile phone voice interaction combined with environmental perception according to claim 2, wherein: Configuring noise suppression for the secondary direction sound signal includes: Using the auxiliary judgment feature as a signal separation criterion, separating the secondary direction sound signal from the noise signal, and determining the power spectrum of the noise signal; determining a power spectrum of the secondary direction sound signal; A filter gain is calculated based on the power spectrum of the noise signal and the power spectrum of the secondary direction sound signal, and noise suppression is performed on the secondary direction sound signal using the filter gain.

5. The method for optimizing mobile phone voice interaction combined with environment perception according to claim 2, characterized in that: Performing sound comparison in different directions on the ambient sound data of each microphone in the ambient sound dataset using the auxiliary judgment feature to determine a main direction sound signal and a secondary direction sound signal, including: Using the auxiliary judgment feature as the main direction sound feature, analyzing the main direction sound feature intensity and noise introduction intensity of the ambient sound data of each microphone, and generating a main sound intensity index and a noise intensity index in each direction; Based on the main sound intensity index and the noise intensity index in each direction, the direction with the maximum main sound intensity and the minimum noise intensity is located, and the main direction sound signal and the secondary direction sound signal are separated from the ambient sound data of each microphone.

6. The method for optimizing mobile phone voice interaction combined with environment perception according to claim 1, characterized in that: Performing primary and secondary direction collaborative noise processing on the primary direction voice data and the secondary direction voice data to generate enhanced voice data includes: Constructing a spectrum model based on the primary direction speech data, repairing the speech spectrum in the secondary direction speech data, and generating a secondary direction speech repair result; Performing noise estimation and restoration on the main direction voice data using the noise spectrum in the secondary direction voice data to generate a main direction voice restoration result; The enhanced speech data is generated using the main direction speech restoration result and the secondary direction speech restoration result.

7. The method for optimizing mobile phone voice interaction combined with environment perception according to claim 1, characterized in that: Collecting the real-time ambient light information of the target user and executing voice interaction parameter constraints in combination with the ambient sound dataset includes: Determine the light and noise characteristics of multiple interactive environment modalities, as well as the corresponding preset interaction parameters; Identifying the real-time ambient light information and the ambient sound dataset using the light characteristics and noise characteristics of the multiple interactive environment modalities to determine a target interactive environment modality; The voice interaction parameter constraints are completed using the preset interaction parameters corresponding to the target interaction environment modality.

8. The method for optimizing mobile phone voice interaction combined with environment perception according to claim 7, characterized in that: Then, based on the user personality voice feature library, the target user enhanced voice data is awakened by a voice interaction module or interactively processed after awakening, including: Training an interactive semantic recognition network based on the user's personalized voice feature library; The interactive semantic recognition network is used to perform semantic recognition on the enhanced voice data of the target user, and the semantic recognition result is used to wake up the voice interaction module or perform interaction processing after wakeup.

9. The method for optimizing mobile phone voice interaction combined with environment perception according to claim 8, characterized in that: If there are multiple rounds of interaction scenarios, the recognition is corrected based on the context of the multiple rounds of interaction scenarios.

10. A mobile phone voice interaction optimization system combined with environmental perception is characterized by: The steps for implementing the method for optimizing mobile phone voice interaction in combination with environment perception according to any one of claims 1 to 9, wherein the mobile phone voice interaction optimization system in combination with environment perception comprises: The ambient sound collection module is used to collect ambient sound through the microphone array of the target mobile phone and generate an ambient sound dataset; a signal suppression analysis module configured with auxiliary judgment features using a target user's personalized voice feature library pre-stored in the voice interaction module, performing comparison and signal suppression analysis on the ambient sound dataset, and generating a microphone signal suppression feature matrix; a noise suppression control module, configured to locate the main direction and the secondary direction based on the microphone signal suppression feature matrix and perform noise suppression control, and then collect the main direction voice data and the secondary direction voice data; a collaborative noise processing module, configured to perform collaborative noise processing in primary and secondary directions using the primary direction voice data and the secondary direction voice data to generate enhanced voice data; The interaction processing module is used to collect the real-time ambient light information of the target user, and execute the voice interaction parameter constraints in combination with the ambient sound data set, and then perform wake-up or post-wake-up interaction processing on the target user's enhanced voice data based on the user's personalized voice feature library.