Intelligent voice interactive Bluetooth earphone operation method and system
By performing multi-modal analysis and sound field separation processing on the operation log and audio data of the intelligent voice interactive Bluetooth headset, combined with time-frequency feature map and environmental fit analysis, the headset's signal loss and control command analysis accuracy problems when processing the original audio data are solved, achieving higher operation accuracy and user experience.
Patent Information
- Application Number
- CN202510624029.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing intelligent voice interactive Bluetooth headsets fail to effectively distinguish user audio from environmental noise when processing original audio data, resulting in signal loss and reduced accuracy of control command analysis, and failure to analyze the conflict between control commands and environment, resulting in conflicts in the headset control results.
By performing multi-modal analysis of the headphone operation log data, the headphone playback task is obtained; the original audio data is separated, the user audio data and environmental audio data are extracted, and the environmental noise band characteristics are analyzed; the headphone interaction intention is analyzed based on the time-frequency feature map; the fit between the environment attributes and the headphone interaction intention is calculated, and the conflict reminder mechanism is constructed to correct the intention when mismatch is not met.
It improves the operation accuracy of intelligent voice interactive Bluetooth headsets, enhances the accuracy of voice recognition, intelligently adjusts the headset working mode to match user intentions, and significantly improves the user experience and the intelligence of the headset.
Smart Images

Figure CN120148495A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an operation method and system for an intelligent voice interactive Bluetooth headset, belonging to the technical field of Bluetooth headset control. Background Art
[0002] An intelligent voice interactive Bluetooth headset is a headset device that combines intelligent voice technology and Bluetooth communication functions. The headset has a voice recognition function, can collect user interaction voice data, perform corresponding control processing on the Bluetooth according to the user's voice, and then realize the voice interaction between the headset and the user.
[0003] The operation method process of the existing intelligent voice interactive Bluetooth headset is as follows: First, turn on the Bluetooth functions of the Bluetooth headset and the paired device (such as a mobile phone, computer), establish a connection through searching and pairing to ensure data transmission. Then, the user wakes up the headset voice interaction function through a specific wake-up word or button, and the microphone immediately captures the sound signal. Then, the headset converts the voice signal into a digital signal and transmits it to the connected device, and the voice recognition engine in the device analyzes and processes it into text. Then, the user's intention is parsed through semantic understanding to clarify the specific operation instruction. Finally, the device executes the operation according to the instruction and feeds back the result to the headset through Bluetooth, and notifies the user that the operation is completed in ways such as voice and indicator light flashing. However, this method does not distinguish the original audio data, and it is easy to cause signal loss during subsequent denoising processing, resulting in a decrease in the accuracy of the final control instruction analysis, and does not analyze the conflict between the current control instruction and the existing environment, thus leading to contradictions in the results of headset control. Therefore, a method that can improve the operation accuracy of the intelligent voice interactive Bluetooth headset is needed. Summary of the Invention
[0004] The present invention provides an operation method and system for an intelligent voice interactive Bluetooth headset, and its main purpose is to improve the operation accuracy of the intelligent voice interactive Bluetooth headset.
[0005] To achieve the above purpose, an operation method for an intelligent voice interactive Bluetooth headset provided by the present invention includes: Obtain the intelligent voice interactive Bluetooth headset to be processed, collect the headset operation log data of the Bluetooth headset, and perform multi-modal analysis processing on the headset operation log data to obtain a headset playback task; Collect the original audio data of the Bluetooth headset about the user, perform sound field separation processing on the original audio data to obtain user audio data and environmental audio data, extract the noise frequency band characteristics corresponding to the environmental audio data, and analyze the environmental attributes corresponding to the Bluetooth headset based on the noise frequency band characteristics; Extract the user audio signal corresponding to the user audio data, extract the time-frequency feature map in the user audio signal, and analyze the headphone interaction intention corresponding to the Bluetooth headset based on the time-frequency feature map; Calculate the fitness between the environmental attribute and the headphone interaction intention. When the fitness is higher than the preset fitness, combine the headphone interaction intention and the headphone playback task, and perform interactive operation processing on the Bluetooth headset to obtain a first operation result; When the fitness is not higher than the preset fitness, use the Bluetooth headset to construct a conflict reminder mechanism corresponding to the environmental attribute, collect the mechanism feedback data of the conflict reminder mechanism, and based on the mechanism feedback data, correct the interaction intention of the headset to obtain a target interaction intention. Combine the target interaction intention and the headphone playback task, and perform interactive operation processing on the Bluetooth headset to obtain a second operation result.
[0006] Optionally, the multi-modal analysis processing of the headphone operation log data to obtain the headphone playback task includes: Clean the headphone operation log data to obtain pure operation log data; Extract the time-series characteristics of the pure operation log data to obtain time-series behavior characteristics; Perform multi-modal fusion on the time-series behavior characteristics to obtain a fused behavior feature vector; Perform behavior pattern clustering processing on the fused behavior feature vector to obtain user behavior pattern labels; Perform task analysis on the user behavior pattern labels to obtain the headphone playback task.
[0007] Optionally, the performing behavior pattern clustering processing on the fused behavior feature vector to obtain user behavior pattern labels includes: Perform vector dimensionality reduction processing on the fused behavior feature vector to obtain a dimensionality-reduced behavior feature vector; Calculate the vector similarity between the dimensionality-reduced behavior feature vectors; Perform clustering processing on the dimensionality-reduced behavior feature vectors based on the vector similarity to obtain clustered behavior feature vectors; Extract the behavior common features in the clustered behavior feature vectors; Perform label annotation processing on the clustered behavior feature vectors based on the behavior common features to obtain user behavior pattern labels.
[0008] Optionally, the performing sound field separation processing on the original audio data to obtain user audio data and environmental audio data includes: Perform audio normalization processing on the original audio data to obtain equal-amplitude standard audio data; Perform spatial spectrum analysis on the equal-amplitude standard audio data to obtain spatial azimuth feature data; Calculate the spatial azimuth tensor corresponding to the spatial azimuth feature data; Based on the spatial azimuth tensor, determine the spatial feature map corresponding to the equal-amplitude standard audio data; Based on the spatial feature map, perform signal beamforming processing on the equal-amplitude standard audio data to obtain a directional enhanced audio signal; Perform sound field separation processing on the directional enhanced audio signal to obtain user audio data and environmental audio data.
[0009] Optionally, the extraction of the noise frequency band characteristics corresponding to the environmental audio data includes: Perform fast Fourier transform processing on the environmental audio data to obtain spectral distribution data; Perform band-pass filtering processing on the spectral distribution data to obtain a target frequency band signal; Perform energy normalization processing on the target frequency band signal to obtain a standard frequency band signal; Use a preset convolutional neural network to perform feature enhancement processing on the standard frequency band signal to obtain an enhanced feature signal; Perform signal feature analysis on the enhanced feature signal to obtain noise frequency band characteristics.
[0010] Optionally, the extraction of the time-frequency feature map in the user audio signal includes: Perform frame splitting processing on the user audio signal to obtain frame-split audio signals, and extract the signal time-frequency diagram corresponding to the frame-split audio signals; Perform multi-scale decomposition processing on the signal time-frequency diagram to obtain a set of time-frequency feature matrices; Perform adaptive enhancement processing on the set of time-frequency feature matrices to obtain an enhanced time-frequency matrix; Perform feature reconstruction processing on the enhanced time-frequency matrix to obtain a reconstructed time-frequency diagram; Perform inverse short-time Fourier transform processing on the reconstructed time-frequency diagram to obtain the time-frequency feature map in the user audio signal.
[0011] Optionally, the calculation of the degree of fit between the environmental attribute and the headphone interaction intention includes: Extract the environmental context features in the environmental attribute, and based on the environmental context features, construct an environmental attribute feature vector corresponding to the environmental attribute; Extract the intention key characters corresponding to the headphone interaction intention, perform semantic parsing on the intention key characters to obtain the interaction intention semantics; Vectorize the interactive intent semantics to obtain an interactive intent semantic vector; Calculate the vector matching index between the environmental attribute feature vector and the interactive intent semantic vector; Calculate the matching degree between the environmental attribute and the headphone interactive intent according to the vector matching index.
[0012] Optionally, the calculating the vector matching index between the environmental attribute feature vector and the interactive intent semantic vector includes: Calculate the vector matching index between the environmental attribute feature vector and the interactive intent semantic vector through the following formula: Where, A represents the vector matching index between the environmental attribute feature vector and the interactive intent semantic vector, B represents the environmental attribute feature vector, D represents the interactive intent semantic vector, represents the norm of the environmental attribute feature vector, represents the norm of the interactive intent semantic vector, represents the Gaussian kernel width.
[0013] Optionally, the constructing the conflict reminder mechanism corresponding to the Bluetooth headset by combining the headphone interactive intent and the environmental attribute includes: Analyze the interactive intent category corresponding to the headphone interactive intent and analyze the environmental attribute category corresponding to the environmental attribute; Analyze the conflict mode between the interactive intent category and the environmental attribute category; Based on the conflict mode, extract the conflict key parameters from the headphone interactive intent and the environmental attribute; Calculate the conflict probability between the conflict key parameters, and set the conflict level between the headphone interactive intent and the environmental attribute according to the conflict probability; Construct the conflict reminder mechanism corresponding to the Bluetooth headset based on the conflict level.
[0014] To solve the above problems, the present invention also provides an intelligent voice interactive Bluetooth headset operating system, and the system includes: A headphone playback task analysis module, configured to obtain an intelligent voice interactive Bluetooth headset to be processed, collect the headphone operation log data of the Bluetooth headset, perform multi-modal analysis processing on the headphone operation log data, and obtain a headphone playback task; An environmental attribute analysis module, configured to collect original audio data of the Bluetooth headset regarding the user, perform sound field separation processing on the original audio data to obtain user audio data and environmental audio data, extract noise frequency band characteristics corresponding to the environmental audio data, and analyze the environmental attributes corresponding to the Bluetooth headset based on the noise frequency band characteristics; A headset interaction intention analysis module, configured to extract a user audio signal corresponding to the user audio data, extract a time-frequency feature map in the user audio signal, and analyze the headset interaction intention corresponding to the Bluetooth headset based on the time-frequency feature map; A first interaction operation processing module, configured to calculate a degree of fit between the environmental attribute and the headset interaction intention, and when the degree of fit is higher than a preset degree of fit, perform interaction operation processing on the Bluetooth headset by combining the headset interaction intention and the headset playback task to obtain a first operation result; A second interaction operation processing module, configured to, when the degree of fit is not higher than the preset degree of fit, use the Bluetooth headset to construct a conflict reminder mechanism corresponding to the environmental attribute, collect mechanism feedback data of the conflict reminder mechanism, correct the interaction intention based on the mechanism feedback data to obtain a target interaction intention, and perform interaction operation processing on the Bluetooth headset by combining the target interaction intention and the headset playback task to obtain a second operation result.
[0015] Compared with the problems described in the background art, the present invention performs multimodal analysis and processing on the headphone operation log data to obtain the headphone playback task, so as to understand the current playback content of the Bluetooth headphone, and further provides a basis for subsequent execution of the interaction operation processing of the Bluetooth headphone. Further, the present invention performs sound field separation processing on the original audio data to obtain user audio data and environmental audio data, and can separate the environmental noise in the original audio data, improving the speech recognition accuracy of the subsequent Bluetooth headphone for the user. The present invention captures the dynamic change characteristics of the speech signal in the time and frequency dimensions by extracting the time-frequency feature map of the user audio signal, effectively separating the speech and noise components, and providing rich and intuitive feature basis for the high-precision recognition of the Bluetooth headphone interaction intention. Further, the present invention can more accurately understand the needs of users when using Bluetooth headphones in different environments by calculating the fit degree between the environmental attributes and the headphone interaction intention, thereby intelligently adjusting the working mode of the headphone or providing an interaction response that better fits the user intention, significantly improving the user experience and the intelligence level of the headphone. Further, it should be understood that when the fit degree is not higher than the preset fit degree, it indicates that the environment does not match the headphone interaction intention. The present invention constructs a conflict reminder mechanism corresponding to the Bluetooth headphone by combining the headphone interaction intention and the environmental attributes, which can timely inform the user that there is a contradiction between the current operation and the environment, avoid invalid operations, and at the same time guide the user to adjust the interaction intention subsequently, making the headphone function match the actual use scenario, and significantly improving the interaction intelligence of the Bluetooth headphone and the convenience and satisfaction of user use. Therefore, the intelligent voice interactive Bluetooth headphone operation method and system provided by the embodiments of the present invention can improve the operation accuracy of the intelligent voice interactive Bluetooth headphone. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic flowchart of the intelligent voice interactive Bluetooth headphone operation method provided by an embodiment of the present invention; Figure 2 It is a schematic diagram of the module for implementing the intelligent voice interactive Bluetooth headphone operation method provided by an embodiment of the present invention.
[0017] The implementation, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0019] An embodiment of the present application provides a method for operating an intelligent voice interactive Bluetooth headset. The execution subject of the method for operating an intelligent voice interactive Bluetooth headset includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiment of the present application. In other words, the method for operating an intelligent voice interactive Bluetooth headset can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.
[0020] Embodiment 1: Refer to Figure 1 As shown, it is a flowchart of a method for operating an intelligent voice interactive Bluetooth headset provided by an embodiment of the present invention. In this embodiment, the method for operating an intelligent voice interactive Bluetooth headset includes: S1. Obtain an intelligent voice interactive Bluetooth headset to be processed, collect the headset operation log data of the Bluetooth headset, and perform multi-modal analysis and processing on the headset operation log data to obtain a headset playback task.
[0021] By performing multi-modal analysis and processing on the headset operation log data of the present invention to obtain a headset playback task, the current playback content of the Bluetooth headset can be understood, and thus a basis is provided for subsequent execution of interactive operation processing on the Bluetooth headset. Among them, the headset operation log data is a collection of information such as user operation behaviors, instruction execution situations, function usage frequencies, and corresponding timestamps recorded during the use of the Bluetooth headset. The headset playback task is a specific audio playback operation that needs to be executed after the Bluetooth headset processes the audio data. Further, the headset operation log data can be collected through the embedded log recording module of the Bluetooth headset, and this module is implemented based on microcontroller programming.
[0022] As an embodiment of the present invention, the multi-modal analysis and processing of the headset operation log data to obtain a headset playback task includes: Performing cleaning processing on the headset operation log data to obtain pure operation log data; Performing time series feature extraction on the pure operation log data to obtain time series behavior features; Performing multi-modal fusion on the time series behavior features to obtain a fused behavior feature vector; Performing behavior pattern clustering processing on the fused behavior feature vector to obtain a user behavior pattern label; Performing task analysis on the user behavior pattern label to obtain a headset playback task.
[0023] Among them, the pure operation log data is the log data that is clean and standardized after the headphone operation log data is cleaned to remove interference information such as duplicate records, error data, and invalid operations; the temporal behavior characteristics are the characteristics reflecting the operation time series law such as operation time intervals, the order of instruction execution, and the change of operation frequency over time obtained after the pure operation log data is subjected to temporal feature extraction; the fused behavior feature vector is the feature vector containing multi-faceted behavior information generated by integrating the temporal behavior characteristics with other modal data (such as user usage scenarios, audio type preferences, etc.) after multi-modal fusion of the temporal behavior characteristics; the user behavior pattern label is the label used to identify the specific behavior pattern of the user assigned to each category after the fused behavior feature vector is subjected to behavior pattern clustering processing and the user behavior is divided into different categories according to the similarity of data characteristics.
[0024] Furthermore, the headphone operation log data can be cleaned through data cleaning algorithms (such as outlier detection and duplicate item deletion) to obtain pure operation log data; the temporal behavior characteristics can be obtained by performing temporal feature extraction on the pure operation log data through a temporal analysis tool (such as the sliding window method); the fused behavior feature vector can be obtained by performing multi-modal fusion on the temporal behavior characteristics through a multi-modal fusion model (such as an LSTM network based on the attention mechanism); the headphone playback task can be obtained by performing task analysis on the user behavior pattern label through a task analysis engine (combining a rule base and a reinforcement learning algorithm).
[0025] Furthermore, as an optional embodiment of the present invention, the process of performing behavior pattern clustering processing on the fused behavior feature vector to obtain the user behavior pattern label includes: Performing vector dimensionality reduction processing on the fused behavior feature vector to obtain a dimensionality-reduced behavior feature vector; Calculating the vector similarity between the dimensionality-reduced behavior feature vectors; Based on the vector similarity, performing clustering processing on the dimensionality-reduced behavior feature vectors to obtain clustered behavior feature vectors; Extracting the behavior common features in the clustered behavior feature vectors; Based on the behavior common features, performing label annotation processing on the clustered behavior feature vectors to obtain the user behavior pattern label.
[0026] Among them, the dimensionality-reduced behavior feature vector is a more concise feature vector that retains the main information obtained after performing vector dimensionality reduction on the fused behavior feature vector. The vector similarity is a quantitative index representing the similarity degree between the dimensionality-reduced behavior feature vectors. The clustered behavior feature vector is a set of vectors with high similarity within each cluster obtained by clustering the dimensionality-reduced behavior feature vectors based on the vector similarity. The behavior common feature is a typical feature that is commonly possessed in the clustered behavior feature vectors and can represent the user behavior patterns included in the clustering, such as common operation habits, similar usage scenarios, etc.
[0027] Furthermore, the fused behavior feature vector can be processed by the principal component analysis (PCA) dimensionality reduction algorithm to obtain the dimensionality-reduced behavior feature vector; the vector similarity between the dimensionality-reduced behavior feature vectors can be calculated by the cosine similarity algorithm; based on the vector similarity, the dimensionality-reduced behavior feature vectors can be clustered by the K-means algorithm to obtain the clustered behavior feature vectors; the behavior common features in the clustered behavior feature vectors can be extracted by the association rule mining algorithm; based on the behavior common features, the clustered behavior feature vectors can be labeled by the decision tree method to obtain the user behavior pattern labels.
[0028] S2. Collect the original audio data of the Bluetooth headset regarding the user, perform sound field separation processing on the original audio data to obtain user audio data and environmental audio data, extract the noise frequency band features corresponding to the environmental audio data, and analyze the environmental attributes corresponding to the Bluetooth headset based on the noise frequency band features.
[0029] In the present invention, by performing sound field separation processing on the original audio data to obtain user audio data and environmental audio data, the environmental noise in the original audio data can be separated, improving the speech recognition accuracy of the subsequent Bluetooth headset regarding the user. Among them, the original audio data is the mixed audio information containing all sound signals such as user speech, environmental noise, and device prompt sounds collected in real time by the Bluetooth headset regarding the user through the built-in microphone; the user audio data is the sound signals directly related to the user, such as user speech and instructions, separated from the original audio data, including speech semantics and intonation features; the environmental audio data is other sound signals in the original audio data except for the user's voice, such as traffic noise and human voices, reflecting the external acoustic environment where the headset is located. Furthermore, the original audio data of the Bluetooth headset regarding the user can be collected by the highly sensitive microphone array built in the Bluetooth headset.
[0030] As an embodiment of the present invention, the processing of performing sound field separation on the original audio data to obtain user audio data and environmental audio data includes: Performing audio normalization processing on the original audio data to obtain equal-amplitude standard audio data; Performing spatial spectrum analysis on the equal-amplitude standard audio data to obtain spatial orientation feature data; Calculating a spatial orientation tensor corresponding to the spatial orientation feature data; Based on the spatial orientation tensor, determining a spatial feature map corresponding to the equal-amplitude standard audio data; Based on the spatial feature map, performing signal beamforming processing on the equal-amplitude standard audio data to obtain a directionally enhanced audio signal; Performing sound field separation processing on the directionally enhanced audio signal to obtain user audio data and environmental audio data.
[0031] Wherein, the equal-amplitude standard audio data is audio data with a unified amplitude specification after audio normalization processing of the original audio data; the spatial orientation feature data is data on the spatial distribution characteristics of the sound signal obtained after spatial spectrum analysis of the equal-amplitude standard audio data; the spatial orientation tensor is a high-order mathematical quantity corresponding to the spatial orientation feature data used to describe the complex characteristics of the sound spatial orientation; the spatial feature map is a visual graph of the sound spatial distribution presented based on the spatial orientation tensor corresponding to the equal-amplitude standard audio data; the directionally enhanced audio signal is audio data with enhanced user direction signals and suppressed noise in other directions after signal beamforming processing of the equal-amplitude standard audio data.
[0032] Furthermore, the original audio data can be subjected to audio normalization processing by performing a normalization operation on the audio amplitude through a digital signal processing algorithm to obtain equal-amplitude standard audio data, such as the least mean square (LMS) algorithm; the equal-amplitude standard audio data can be subjected to spatial spectrum analysis through a fast Fourier transform combined with a spatial spectrum estimation algorithm to obtain spatial orientation feature data; the spatial orientation tensor corresponding to the spatial orientation feature data can be calculated through a tensor calculation model and eigenvector operations. First, the spatial orientation feature data is represented as the sum of multiple low-rank tensors, and the spatial orientation feature data is mapped to a low-dimensional subspace to obtain eigenvectors. Then, an outer product operation is performed on these eigenvectors, and then iterative optimization is carried out according to the least squares method to continuously adjust each factor matrix to minimize the reconstruction error, and finally the spatial orientation tensor is calculated; based on the spatial orientation tensor, the spatial feature map corresponding to the equal-amplitude standard audio data can be determined through a data visualization algorithm and a spatial mapping model; based on the spatial feature map, the equal-amplitude standard audio data can be subjected to signal beamforming processing through an adaptive beamforming algorithm to obtain a directional enhanced audio signal; the directional enhanced audio signal can be subjected to sound field separation processing through a blind source separation algorithm to obtain user audio data and environmental audio data.
[0033] By extracting the noise frequency band characteristics corresponding to the environmental audio data in the present invention, the frequency distribution characteristics, energy concentration regions, noise dominant components, and types of environmental noise of the environmental audio data can be understood. Based on the noise frequency band characteristics, the environmental attributes corresponding to the Bluetooth headset can be analyzed, and then the environmental characteristics of the Bluetooth headset can be understood, which is convenient for subsequent execution of interaction operation processing on the Bluetooth headset. Among them, the noise frequency band characteristics are characteristic information such as the energy distribution, spectral characteristics, and prominent manifestations of specific noise frequency components corresponding to the environmental audio data in different frequency intervals, and the environmental attributes are characteristic descriptions of the usage environment corresponding to the Bluetooth headset, including but not limited to the noise level, noise type, spatial distribution characteristics of sound, etc. Further, based on the noise frequency band characteristics, the environmental attributes corresponding to the Bluetooth headset are analyzed by comparing with a typical environmental noise database. For example, the energy distribution and characteristic frequencies of the noise frequency band are matched with the standard spectral characteristics of scenarios such as traffic noise, industrial noise, and human voices in the database, and the current environment is judged to be a crowded area such as a subway car, a factory workshop, or a shopping mall according to the matching degree.
[0034] As an embodiment of the present invention, the extraction of the noise frequency band characteristics corresponding to the environmental audio data includes: Performing a fast Fourier transform process on the environmental audio data to obtain spectral distribution data; Performing a band-pass filtering process on the spectral distribution data to obtain a target frequency band signal; Perform energy normalization processing on the target frequency band signal to obtain a standard frequency band signal; Use a preset convolutional neural network to perform feature enhancement processing on the standard frequency band signal to obtain an enhanced feature signal; Perform signal feature analysis on the enhanced feature signal to obtain noise frequency band features.
[0035] Among them, the spectrum distribution data is the frequency and amplitude distribution data obtained by converting the time-domain signal into the frequency domain after performing fast Fourier transform processing on the environmental audio data; the target frequency band signal is the signal obtained by filtering out a specific frequency range after performing band-pass filtering on the spectrum distribution data; the standard frequency band signal is the data obtained by performing energy normalization processing on the target frequency band signal to make the signal energy in a unified scale range; the preset convolutional neural network is a deep learning model that has been pre-trained with a large amount of environmental noise data and is used to automatically extract signal features; the enhanced feature signal is the noise feature data that is more discriminative and robust mined after using the preset convolutional neural network to perform feature enhancement processing on the standard frequency band signal.
[0036] Furthermore, the fast Fourier transform processing can be performed on the environmental audio data through the FFT operation module built in the digital signal processing chip to obtain the spectrum distribution data; the band-pass filtering processing can be performed on the spectrum distribution data through the programmable digital filter to obtain the target frequency band signal; the energy normalization processing can be performed on the target frequency band signal through the normalization algorithm module to obtain the standard frequency band signal; use the preset convolutional neural network to perform feature enhancement processing on the standard frequency band signal to obtain the enhanced feature signal, and the steps are as follows: First, input the standard frequency band signal into the input layer of the preset convolutional neural network, and the input layer will adjust the signal format to adapt to the subsequent network layer processing; then the signal performs convolutional operations through multiple convolutional kernels in the convolutional layer to extract the local features of the signal, and then passes through the pooling layer to perform downsampling on the feature map to reduce the data volume and enhance the robustness of the features; finally, the features are integrated through the fully connected layer to output the enhanced feature signal; the signal feature analysis can be performed on the enhanced feature signal through a feature extraction algorithm based on machine learning to obtain the noise frequency band features, such as the independent component analysis algorithm.
[0037] S3. Extract the user audio signal corresponding to the user audio data, extract the time-frequency feature map in the user audio signal, and analyze the headphone interaction intention corresponding to the Bluetooth headset based on the time-frequency feature map.
[0038] By extracting the time-frequency feature map from the user audio signal, the dynamic change features of the speech signal in the time and frequency dimensions can be captured, effectively separating the speech and noise components, and providing rich and intuitive feature basis for the high-precision recognition of the interaction intention of the Bluetooth headset. Among them, the time-frequency feature map is a visual feature expression that converts the time-domain audio signal in the user audio signal into a two-dimensional space of time-frequency, intuitively presenting the change of the signal frequency component over time. Further, the user audio signal corresponding to the user audio data can be extracted through audio decoding technology.
[0039] As an embodiment of the present invention, the extraction of the time-frequency feature map from the user audio signal includes: Perform frame segmentation processing on the user audio signal to obtain frame-segmented audio signals, and extract the signal time-frequency map corresponding to the frame-segmented audio signals; Perform multi-scale decomposition processing on the signal time-frequency map to obtain a set of time-frequency feature matrices; Perform adaptive enhancement processing on the set of time-frequency feature matrices to obtain an enhanced time-frequency matrix; Perform feature reconstruction processing on the enhanced time-frequency matrix to obtain a reconstructed time-frequency map; Perform inverse short-time Fourier transform processing on the reconstructed time-frequency map to obtain the time-frequency feature map in the user audio signal.
[0040] Among them, the frame-segmented audio signal is a plurality of short-time segment audio clips obtained by performing frame segmentation processing on the user audio signal. The signal time-frequency map is a two-dimensional visualization map that converts the time-domain signal of the frame-segmented audio signal into a frequency change over time. The set of time-frequency feature matrices is a combination of multiple feature matrices extracted at different scales after performing multi-scale decomposition processing on the signal time-frequency map. The enhanced time-frequency matrix is an optimized matrix that strengthens the key features of the speech and suppresses the noise interference after performing adaptive enhancement processing on the set of time-frequency feature matrices. The reconstructed time-frequency map is a time-frequency domain map that optimizes the details and improves the quality through technologies such as generative adversarial networks after performing feature reconstruction processing on the enhanced time-frequency matrix.
[0041] Further, the user audio signal can be frame-processed through a sliding window algorithm to obtain a framed audio signal. The signal time-frequency diagram corresponding to the framed audio signal can be extracted through a short-time Fourier transform (STFT) algorithm. The signal time-frequency diagram can be multi-scale decomposed through a wavelet transform algorithm to obtain a set of time-frequency feature matrices. The set of time-frequency feature matrices can be adaptively enhanced through an adaptive gain control algorithm based on deep learning to obtain an enhanced time-frequency matrix. The enhanced time-frequency matrix can be feature-reconstructed through a generative adversarial network (GAN) model to obtain a reconstructed time-frequency diagram. The reconstructed time-frequency diagram can be inverse short-time Fourier transformed through an inverse short-time Fourier transform (iSTFT) algorithm to obtain the time-frequency feature map in the user audio signal.
[0042] Based on the time-frequency feature map, the present invention analyzes the headset interaction intention corresponding to the Bluetooth headset, and can understand the purpose that the user expects to achieve for related operations on the Bluetooth headset. Among them, the headset interaction intention is the target or behavior direction that the user corresponding to the Bluetooth headset expects to achieve by operating or issuing commands through the headset. Further, the time-frequency feature map is matched with a predefined headset interaction intention pattern library to analyze the headset interaction intention corresponding to the Bluetooth headset. The headset interaction intention pattern library contains typical time-frequency feature patterns corresponding to various known headset interaction intentions. For example, the intention of playing music may correspond to a pattern with concentrated energy in a specific frequency range and relatively stable frequency changes, while the intention of adjusting the volume may be manifested as an obvious upward or downward trend in frequency within a short period of time. By comparing the feature vectors with the patterns in the pattern library, the most matching pattern is found, thereby preliminarily determining the headset interaction intention.
[0043] S4. Calculate the fit degree between the environmental attribute and the headset interaction intention. When the fit degree is higher than the preset fit degree, the interaction operation processing of the Bluetooth headset is performed by combining the headset interaction intention and the headset playback task to obtain a first operation result.
[0044] By calculating the fit degree between the environmental attribute and the headset interaction intention, the present invention can more accurately understand the needs of the user when using the Bluetooth headset in different environments, thereby intelligently adjusting the working mode of the headset or providing an interaction response that better fits the user's intention, significantly improving the user experience and the intelligence level of the headset. Among them, the fit degree represents a numerical value of the degree of tight association between the environmental attribute and the headset interaction intention.
[0045] As an embodiment of the present invention, the calculation of the fit degree between the environmental attribute and the headset interaction intention includes: Extract the environmental context features in the environmental attributes, and construct an environmental attribute feature vector corresponding to the environmental attributes based on the environmental context features; Extract the key intent characters corresponding to the headphone interaction intent, and perform semantic parsing on the key intent characters to obtain the interaction intent semantics; Perform vectorization processing on the interaction intent semantics to obtain an interaction intent semantic vector; Calculate the vector matching index between the environmental attribute feature vector and the interaction intent semantic vector; Calculate the matching degree between the environmental attributes and the headphone interaction intent according to the vector matching index.
[0046] Among them, the environmental context features are specific feature information in the environmental attributes that can reflect the acoustic, physical, and location conditions of the environment. The environmental attribute feature vector is a set of environmental attribute-related information represented in vector form after feature extraction and quantization corresponding to the environmental attributes. The key intent characters are the core characters or words corresponding to the headphone interaction intent that are used to represent and identify specific headphone interaction intents. The interaction intent semantics is the meaning and explanatory content of the headphone interaction intent obtained after semantic parsing of the key intent characters. The interaction intent semantic vector is a vector form that contains interaction intent semantic information and is represented in a vector space after vectorization processing of the interaction intent semantics. The vector matching index represents a numerical index calculated by a certain measurement method between the environmental attribute feature vector and the interaction intent semantic vector, which reflects the matching degree or similarity between the two vectors at the feature and semantic levels.
[0047] Furthermore, the environmental context features in the environmental attributes can be extracted by an acoustic feature extraction method. For example, for the acoustic features in the environment, such as noise intensity and frequency distribution, audio processing technology can be used to achieve this. Based on the environmental context features, an environmental attribute feature vector corresponding to the environmental attributes can be constructed through a feature selection and normalization algorithm (such as max-min normalization). The key intent characters corresponding to the headphone interaction intent can be extracted by a keyword extraction algorithm in natural language processing (such as TF-IDF, TextRank). The key intent characters can be semantically parsed by a semantic parsing model (such as dependency syntax analysis, knowledge graph reasoning) to obtain the interaction intent semantics. The interaction intent semantics can be vectorized by a word vector model (such as Word2Vec, BERT) to obtain an interaction intent semantic vector. The vector matching index is converted into a percentage form to obtain the matching degree between the environmental attributes and the headphone interaction intent.
[0048] Further, as an alternative embodiment of the present invention, calculating the vector matching index between the environmental attribute feature vector and the interaction intention semantic vector includes: Calculating the vector matching index between the environmental attribute feature vector and the interaction intention semantic vector through the following formula: where A represents the vector matching index between the environmental attribute feature vector and the interaction intention semantic vector, B represents the environmental attribute feature vector, and D represents the interaction intention semantic vector. represents the norm of the environmental attribute feature vector. represents the norm of the interaction intention semantic vector. represents the Gaussian kernel width.
[0049] where is the dimension of the environmental attribute feature vector, that is, the number of elements of the environmental attribute feature vector.
[0050] It should be understood that when the matching degree is higher than the preset matching degree, it indicates that the environment is highly compatible with the headphone interaction intention. Then, the present invention combines the headphone interaction intention and the headphone playback task, performs an interaction operation process on the Bluetooth headphone, and obtains a first operation result, which can efficiently meet the user's usage requirements in the current environment and greatly improve the intelligent experience of the Bluetooth headphone and the convenience of user operation. Among them, the preset matching degree is a standard value preset for judging whether the matching degree between the environmental attribute and the headphone interaction intention reaches the standard. The first operation result is the specific execution effect generated after performing the interaction operation process on the Bluetooth headphone by combining the headphone interaction intention and the headphone playback task. Further, by combining the headphone interaction intention and the headphone playback task, the corresponding hardware function modules (such as the audio decoding module, Bluetooth communication module, etc.) can be called through the control chip and program algorithm built in the headphone to perform the interaction operation process on the Bluetooth headphone and obtain the first operation result.
[0051] S5. When the matching degree is not higher than the preset matching degree, combine the headphone interaction intention and the environmental attribute, construct a conflict reminder mechanism corresponding to the Bluetooth headphone, collect the mechanism feedback data of the conflict reminder mechanism, and based on the mechanism feedback data, correct the interaction intention to obtain a target interaction intention. Then, combine the target interaction intention and the headphone playback task, perform an interaction operation process on the Bluetooth headphone, and obtain a second operation result.
[0052] It should be understood that when the matching degree is not higher than the preset matching degree, it indicates that the interaction intention between the environment and the earphone is not suitable. By combining the earphone interaction intention and the environmental attributes, the present invention constructs a conflict reminder mechanism corresponding to the Bluetooth earphone, which can timely inform the user that there is a contradiction between the current operation and the environment, avoid ineffective operations, and at the same time guide the user to adjust the interaction intention subsequently, so that the earphone function matches the actual use scenario, significantly improving the interaction intelligence of the Bluetooth earphone and the convenience and satisfaction of user use. Among them, the conflict reminder mechanism is to combine the earphone interaction intention and the environmental attributes to construct a reminder strategy corresponding to the Bluetooth earphone for prompting the user that there is a contradiction or mismatch between the current earphone operation intention and the use environment.
[0053] As an embodiment of the present invention, the construction of the conflict reminder mechanism corresponding to the Bluetooth earphone by combining the earphone interaction intention and the environmental attributes includes: Analyze the interaction intention category corresponding to the earphone interaction intention, and analyze the environmental attribute category corresponding to the environmental attributes; Analyze the conflict mode between the interaction intention category and the environmental attribute category; Based on the conflict mode, extract the conflict key parameters from the earphone interaction intention and the environmental attributes; Calculate the conflict probability between the conflict key parameters, and set the conflict level between the earphone interaction intention and the environmental attributes according to the conflict probability; Based on the conflict level, construct the conflict reminder mechanism corresponding to the Bluetooth earphone.
[0054] Among them, the interaction intention category is the classification of the function operation type corresponding to the earphone interaction intention, the environmental attribute category is the classification of the scene characteristics corresponding to the environmental attributes, the conflict mode is the common contradiction manifestation form between the interaction intention category and the environmental attribute category, the conflict key parameter is the core element that plays a decisive role in the conflict determination in the earphone interaction intention and the environmental attributes, the conflict probability is the quantification value of the mismatch degree between the conflict key parameters, and the conflict level is the grading identifier of the conflict severity between the earphone interaction intention and the environmental attributes.
[0055] Furthermore, the interaction intention category corresponding to the headphone interaction intention can be analyzed through natural language processing and intention recognition algorithms; the environmental attribute category corresponding to the environmental attributes can be analyzed through an environmental sensor data classification model; the conflict pattern between the interaction intention category and the environmental attribute category can be analyzed through historical conflict data mining and clustering algorithms. For example, the DBSCAN algorithm is used to cluster combinations such as "playing low-volume music (interaction intention category)" and "noisy street (environmental attribute category)" in historical data to identify common conflict patterns where the content cannot be clearly heard due to the mismatch between volume and environmental noise. Based on the conflict pattern, conflict key parameters can be extracted from the headphone interaction intention and the environmental attributes through expert experience rules. For example, experts determine headphone volume, environmental noise decibel value, audio type, etc. as key parameters according to the above conflict pattern. When the environmental noise exceeds 70 decibels and the headphone volume is lower than 40%, conflicts are likely to occur. The conflict probability between the conflict key parameters can be calculated through a probability statistical model. For example, using the Bayesian probability model and combining the frequency of conflicts when the environmental noise is between 60 - 80 decibels and the headphone volume is lower than 30% in historical data, the conflict occurrence probability of similar parameter combinations in the current scenario is calculated. According to the conflict probability, the conflict level between the headphone interaction intention and the environmental attributes is set through a preset probability-level mapping rule. For example, when the conflict probability is higher than 80%, it is set as the "severe conflict" level, when the probability is between 50% - 80%, it is set as the "moderate conflict" level, and when it is lower than 50%, it is set as the "mild conflict" level. Based on the conflict level, a conflict reminder mechanism corresponding to the Bluetooth headset is constructed by matching a predefined reminder strategy library and a dynamic adjustment mechanism. For example, for the "severe conflict" level, the strategy of "high-volume voice alarm + mobile phone pop-up window prompt to adjust volume" is retrieved from the strategy library, and the reminder frequency and content are dynamically adjusted according to the user's real-time feedback.
[0056] By feeding back data based on the said mechanism to correct the interaction intention of the earphone, the present invention can accurately identify and correct the deviation between the original intention and the actual environment and the real needs of the user. Combining the target interaction intention and the earphone playback task, it performs interaction operation processing on the Bluetooth earphone to obtain a second operation result, making the earphone function execution more in line with the actual usage scenario, effectively avoiding misoperations, and significantly improving the accuracy of the user's control experience and device response. Among them, the mechanism feedback data is a set of user response information of the conflict reminder mechanism, the target interaction intention is an operation instruction that is more in line with the actual needs obtained by correcting the earphone interaction intention based on the mechanism feedback data, and the second operation result is the actual function execution effect generated after performing interaction operation processing on the Bluetooth earphone by combining the target interaction intention and the earphone playback task. Further, the mechanism feedback data of the conflict reminder mechanism can be collected by deploying an event listening module and a log recording system on the earphone and the terminal device connected to it; based on the mechanism feedback data, natural language processing technology is used to parse the semantic meaning of the user feedback, and combined with machine learning algorithms to analyze the historical interaction pattern to correct the earphone interaction intention to obtain the target interaction intention. For example, when the mechanism feedback data contains the user's voice feedback "This music is too soft outside and can't be heard clearly", using the sentiment analysis and semantic understanding technology of natural language processing, it is identified that the user is dissatisfied with the current music volume and hopes to increase the volume. At the same time, combined with machine learning algorithms to analyze the user's historical adjustment pattern of music volume in the outdoor environment, it is found that the user usually increases the volume to 70%-80% to meet the listening needs. Based on these analyses, the original earphone interaction intention of "playing music at 30% volume" is corrected to "playing music at 75% volume", thus obtaining the target interaction intention; by combining the target interaction intention and the earphone playback task, the interaction operation processing on the Bluetooth earphone can be performed by calling the built-in control program and hardware driver of the earphone according to the preset task scheduling rules to obtain the second operation result.
[0057] Compared with the problems described in the background art, the present invention performs multi-modal analysis and processing on the headphone operation log data to obtain the headphone playback task, so as to understand the current playback content of the Bluetooth headphone, and further provides a basis for subsequent execution of the interaction operation processing on the Bluetooth headphone. Further, the present invention performs sound field separation processing on the original audio data to obtain user audio data and environmental audio data, and can separate the environmental noise in the original audio data, improving the speech recognition accuracy of the subsequent Bluetooth headphone regarding the user. The present invention captures the dynamic change characteristics of the speech signal in the time and frequency dimensions by extracting the time-frequency feature map of the user audio signal, effectively separating the speech and noise components, and providing rich and intuitive feature basis for the high-precision recognition of the Bluetooth headphone interaction intention. Further, the present invention can more accurately understand the needs of users when using Bluetooth headphones in different environments by calculating the fit degree between the environmental attributes and the headphone interaction intention, so as to intelligently adjust the working mode of the headphone or provide an interaction response that better fits the user intention, significantly improving the user experience and the intelligence level of the headphone. Further, it should be understood that when the fit degree is not higher than the preset fit degree, it indicates that the environment does not match the headphone interaction intention. The present invention constructs a conflict reminder mechanism corresponding to the Bluetooth headphone by combining the headphone interaction intention and the environmental attributes, which can timely inform the user that there is a conflict between the current operation and the environment, avoid ineffective operations, and at the same time guide the user to adjust the interaction intention subsequently, making the headphone function match the actual use scenario, significantly improving the interaction intelligence of the Bluetooth headphone and the convenience and satisfaction of user use. Therefore, the intelligent voice interactive Bluetooth headphone operation method and system provided by the embodiments of the present invention can improve the operation accuracy of the intelligent voice interactive Bluetooth headphone.
[0058] Embodiment 2: As Figure 2 shown, it is a functional module diagram of an intelligent voice interactive Bluetooth headphone operating system of the present invention.
[0059] The intelligent voice interactive Bluetooth headphone operating system 200 of the present invention can be installed in an electronic device. According to the functions achieved, the intelligent voice interactive Bluetooth headphone operating system can include a headphone playback task analysis module 201, an environmental attribute analysis module 202, a headphone interaction intention analysis module 203, a first interaction operation processing module 204, and a second interaction operation processing module 205. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by an electronic device processor and can complete fixed functions, and are stored in the memory of the electronic device.
[0060] In the embodiments of the present invention, the functions of each module / unit are as follows: The headphone playback task analysis module 201 is used to obtain an intelligent voice interactive Bluetooth headphone to be processed, collect headphone operation log data of the Bluetooth headphone, perform multi-modal analysis and processing on the headphone operation log data, and obtain a headphone playback task; The environmental attribute analysis module 202 is used to collect original audio data of the Bluetooth headphone about the user, perform sound field separation processing on the original audio data to obtain user audio data and environmental audio data, extract noise frequency band characteristics corresponding to the environmental audio data, and analyze the environmental attributes corresponding to the Bluetooth headphone based on the noise frequency band characteristics; The headphone interaction intention analysis module 203 is used to extract a user audio signal corresponding to the user audio data, extract a time-frequency feature map in the user audio signal, and analyze the headphone interaction intention corresponding to the Bluetooth headphone based on the time-frequency feature map; The first interaction operation processing module 204 is used to calculate the degree of fit between the environmental attribute and the headphone interaction intention. When the degree of fit is higher than a preset degree of fit, the interaction operation processing of the Bluetooth headphone is performed by combining the headphone interaction intention and the headphone playback task to obtain a first operation result; The second interaction operation processing module 205 is used to, when the degree of fit is not higher than the preset degree of fit, use the Bluetooth headphone to construct a conflict reminder mechanism corresponding to the environmental attribute, collect mechanism feedback data of the conflict reminder mechanism, correct the interaction intention based on the mechanism feedback data to obtain a target interaction intention, and perform the interaction operation processing of the Bluetooth headphone by combining the target interaction intention and the headphone playback task to obtain a second operation result.
[0061] Specifically, each module in the intelligent voice interactive Bluetooth headphone operating system 200 in the embodiments of the present invention adopts the same technical means as those in the Figure 1 intelligent voice interactive Bluetooth headphone operation method described above, and can produce the same technical effects, which will not be elaborated here.
[0062] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.
[0063] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for operating an intelligent voice interactive Bluetooth headset, characterized in that: The method comprises: Acquire the intelligent voice interactive Bluetooth headset to be processed, collect the headset operation log data of the Bluetooth headset, perform multimodal analysis on the headset operation log data, and obtain the headset playback task; Collecting original audio data of the Bluetooth headset about the user, performing sound field separation processing on the original audio data to obtain user audio data and environmental audio data, extracting noise frequency band characteristics corresponding to the environmental audio data, and analyzing the environmental attributes corresponding to the Bluetooth headset based on the noise frequency band characteristics; Extracting a user audio signal corresponding to the user audio data, extracting a time-frequency feature spectrum in the user audio signal, and analyzing an earphone interaction intention corresponding to the Bluetooth earphone based on the time-frequency feature spectrum; Calculating the degree of fit between the environmental attribute and the headset interaction intention, and when the degree of fit is higher than a preset degree of fit, combining the headset interaction intention and the headset playback task, performing an interactive operation process on the Bluetooth headset to obtain a first operation result; When the degree of fit is not higher than a preset degree of fit, the Bluetooth headset is used to construct a conflict reminder mechanism corresponding to the environmental attributes, and mechanism feedback data of the conflict reminder mechanism is collected. Based on the mechanism feedback data, the headset interaction intention is corrected to obtain the target interaction intention. In combination with the target interaction intention and the headset playback task, the interaction operation processing of the Bluetooth headset is performed to obtain a second operation result.
2. The intelligent voice interactive Bluetooth headset operation method according to claim 1, characterized in that: The performing multimodal analysis and processing on the headphone operation log data to obtain the headphone playback task includes: Cleaning the headphone operation log data to obtain pure operation log data; Extracting time series features from the clean operation log data to obtain time series behavior features; Performing multi-modal fusion on the temporal behavior features to obtain a fused behavior feature vector; Performing behavior pattern clustering processing on the fused behavior feature vector to obtain a user behavior pattern label; Perform task analysis on the user behavior pattern label to obtain a headphone playback task.
3. The intelligent voice interactive Bluetooth headset operation method according to claim 2, characterized in that: The performing behavior pattern clustering processing on the fused behavior feature vector to obtain a user behavior pattern label includes: Performing vector dimensionality reduction processing on the fused behavior feature vector to obtain a reduced-dimensional behavior feature vector; Calculating the vector similarity between the dimensionality reduction behavior feature vectors; Based on the vector similarity, clustering the dimension-reduced behavior feature vector to obtain a clustered behavior feature vector; Extracting common behavior features from the clustered behavior feature vector; Based on the common characteristics of the behaviors, the clustered behavior feature vectors are labeled to obtain user behavior pattern labels.
4. The intelligent voice interactive Bluetooth headset operation method according to claim 1, characterized in that: The performing sound field separation processing on the original audio data to obtain user audio data and environmental audio data includes: Performing audio normalization processing on the original audio data to obtain equal-amplitude standard audio data; Performing spatial spectrum analysis on the equal-amplitude standard audio data to obtain spatial orientation feature data; Calculating the spatial orientation tensor corresponding to the spatial orientation feature data; Based on the spatial orientation tensor, determining a spatial feature spectrum corresponding to the equal-amplitude standard audio data; Based on the spatial feature map, signal beamforming processing is performed on the equal-amplitude standard audio data to obtain a directional enhanced audio signal; The directional enhanced audio signal is subjected to sound field separation processing to obtain user audio data and environmental audio data.
5. The intelligent voice interactive Bluetooth headset operation method according to claim 1, characterized in that: The extracting the noise frequency band features corresponding to the environmental audio data includes: Performing fast Fourier transform processing on the environmental audio data to obtain frequency spectrum distribution data; Performing bandpass filtering on the spectrum distribution data to obtain a target frequency band signal; Performing energy standardization processing on the target frequency band signal to obtain a standard frequency band signal; Using a preset convolutional neural network to perform feature enhancement processing on the standard frequency band signal to obtain an enhanced feature signal; Signal characteristic analysis is performed on the enhanced characteristic signal to obtain noise frequency band characteristics.
6. The intelligent voice interactive Bluetooth headset operation method according to claim 1, characterized in that: The extracting the time-frequency feature spectrum from the user audio signal includes: Performing frame processing on the user audio signal to obtain a framed audio signal, and extracting a signal time-frequency diagram corresponding to the framed audio signal; Performing multi-scale decomposition processing on the signal time-frequency diagram to obtain a set of time-frequency feature matrices; Performing adaptive enhancement processing on the time-frequency feature matrix set to obtain an enhanced time-frequency matrix; Performing feature reconstruction processing on the enhanced time-frequency matrix to obtain a reconstructed time-frequency graph; The reconstructed time-frequency graph is subjected to inverse short-time Fourier transform processing to obtain a time-frequency feature graph in the user audio signal.
7. The intelligent voice interactive Bluetooth headset operation method according to claim 1, characterized in that: The calculating the degree of compatibility between the environmental attribute and the headset interaction intention includes: Extracting environmental context features from the environmental attributes, and constructing an environmental attribute feature vector corresponding to the environmental attributes based on the environmental context features; Extracting the intent key characters corresponding to the headset interaction intent, performing semantic analysis on the intent key characters, and obtaining the interaction intent semantics; Vectorizing the interaction intention semantics to obtain an interaction intention semantic vector; Calculating a vector fit index between the environmental attribute feature vector and the interaction intention semantic vector; The degree of fit between the environmental attribute and the headset interaction intention is calculated according to the vector fit index.
8. The intelligent voice interactive Bluetooth headset operation method according to claim 7, characterized in that: The calculating of the vector fit index between the environmental attribute feature vector and the interaction intention semantic vector includes: The vector fit index between the environmental attribute feature vector and the interaction intention semantic vector is calculated by the following formula: Among them, A represents the vector fit index between the environmental attribute feature vector and the interaction intention semantic vector, B represents the environmental attribute feature vector, and D represents the interaction intention semantic vector. represents the norm of the environmental attribute feature vector, represents the norm of the interaction intention semantic vector, represents the Gaussian kernel width.
9. The intelligent voice interactive Bluetooth headset operation method according to claim 1, characterized in that: The step of combining the headset interaction intention and the environmental attribute to construct a conflict reminder mechanism corresponding to the Bluetooth headset includes: analyzing the interaction intention category corresponding to the headset interaction intention, and analyzing the environmental attribute category corresponding to the environmental attribute; analyzing conflict patterns between the interaction intention categories and the environmental attribute categories; Based on the conflict mode, extracting conflict key parameters from the headset interaction intention and the environmental attributes; Calculating the conflict probability between the conflict key parameters, and setting the conflict level between the headset interaction intention and the environmental attribute according to the conflict probability; Based on the conflict level, a conflict reminder mechanism corresponding to the Bluetooth headset is constructed.
10. An intelligent voice interactive Bluetooth headset operating system, characterized in that: The system comprises: The headphone playback task analysis module is used to obtain the intelligent voice interactive Bluetooth headset to be processed, collect the headphone operation log data of the Bluetooth headset, perform multimodal analysis on the headphone operation log data, and obtain the headphone playback task; An environmental attribute analysis module is used to collect the original audio data of the Bluetooth headset about the user, perform sound field separation processing on the original audio data to obtain user audio data and environmental audio data, extract the noise frequency band characteristics corresponding to the environmental audio data, and analyze the environmental attributes corresponding to the Bluetooth headset based on the noise frequency band characteristics; A headset interaction intention analysis module, used to extract a user audio signal corresponding to the user audio data, extract a time-frequency feature spectrum in the user audio signal, and analyze the headset interaction intention corresponding to the Bluetooth headset based on the time-frequency feature spectrum; A first interactive operation processing module, used for calculating the degree of fit between the environmental attribute and the headset interaction intention, and when the degree of fit is higher than a preset degree of fit, combining the headset interaction intention and the headset playback task, performing interactive operation processing on the Bluetooth headset to obtain a first operation result; The second interactive operation processing module is used to use the Bluetooth headset to construct a conflict reminder mechanism corresponding to the environmental attribute when the degree of fit is not higher than a preset degree of fit, and collect mechanism feedback data of the conflict reminder mechanism, and based on the mechanism feedback data, perform intention correction on the headset interaction intention to obtain the target interaction intention, and combine the target interaction intention and the headset playback task to perform interactive operation processing on the Bluetooth headset to obtain a second operation result.