Earphone artificial intelligence interaction method, system, storage medium and earphone
Patent Information
- Application Number
- CN202510951032.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-07-10
AI Technical Summary
[0003]目前,基于耳机的智能交互方法,通常是通过对数据库中存储的语音数据进行语音特征提取,利用统计学的方式进行建模,构造声学特征,将声学特征和文本转换进行融合方式实现交互语音的生成,然而,该方式过于依赖文本数据的分词和语音数据的特征筛选,从而导致模型在使用时受到限制,并且,该方式所形成的交互语音的质量也会受到数据特征筛选的影响,从而导致质量不高、清晰度不够
[0017]本发明当中的耳机的人工智能交互方法、系统、存储介质及耳机,通过采集耳机的当前状态,基于当前状态控制耳机进入到对应的控制模式,并在对应的控制模式下,通过采集用户的当前图像来判断耳机所使用的用户是否为常用用户,根据耳机在对话模式下的语音信号进行预处理,以使语音信号的频谱平坦化,提高短时信号的连续性;将语音信号所得到的语音特征输入至对应的语音情感识别模型中,将深度频域特征向量和深度时序特征向量进行融合,以实现特征间的信息交互,提高信息处理的性能,利用语言情感模型实现语音特征中语音时间特征和语音频率特征进行融合,以捕捉时间特征和频率特征之间的互补信息,减少信息冗余,提高模型的识别效果和识别精度,利用所识别出的情感类型得到对应的语音模板,利用语音特征和语音模板生成交互语音,以实现耳机与用户之间的语音交互。
Smart Images

Figure CN120872144B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to an artificial intelligence interaction method, system, storage medium, and earphone for headphones. Background Technology
[0002] With the rapid development of technology and the improvement of people's living standards, voice devices have become common devices used for entertainment and leisure in people's lives.
[0003] Currently, headphone-based intelligent interaction methods typically extract speech features from speech data stored in a database, use statistical methods to model and construct acoustic features, and then fuse these acoustic features with text conversion to generate interactive speech. However, this method relies too heavily on word segmentation of text data and feature selection of speech data, which limits the model's usability. Furthermore, the quality of the interactive speech generated by this method is also affected by the data feature selection, resulting in low quality and insufficient clarity. Summary of the Invention
[0004] Based on this, the purpose of the present invention is to provide an artificial intelligence interaction method, system, storage medium, and earphone for headphones, so as to at least solve the shortcomings of the above-mentioned technologies.
[0005] This invention proposes an artificial intelligence interaction method for headphones, wherein the headphones are equipped with at least an image acquisition device and a voice acquisition device, and the method includes: Obtain the current state of the headphones, and control the headphones to enter the corresponding control mode based on the current state; In the control mode, the image acquisition device captures the user's current image in real time, and controls the earphone to enter the corresponding dialogue mode based on the current image; In the dialogue mode, the user's voice signal is collected in real time by the voice acquisition device, and the voice signal is preprocessed to obtain the corresponding voice features; A language emotion recognition model is constructed, and the speech features are input into the speech emotion recognition model to identify the emotion type in the speech signal; According to the emotion type, the corresponding voice template is obtained from the preset template database, and the corresponding interactive voice is generated based on the voice template and the voice features. The interactive voice is then used to realize the voice interaction between the earphone and the user.
[0006] Furthermore, the current state includes a wearing state and an idle state, the control mode includes a noise cancellation mode, an ambient mode, and an interaction mode, and the step of obtaining the current state of the headphones and controlling the headphones to enter the corresponding control mode based on the current state includes: When the headphones are currently in a wearing state, the headphones are controlled to enter a noise cancellation mode. In the noise cancellation mode, the headphones will pulse suppress the noise signal of the surrounding environment. When a first control signal is received from the headphones in the noise cancellation mode, the headphones are controlled to enter the ambient mode according to the first control signal. When a second control signal is received from the headphones in the noise cancellation mode, the headphones are controlled to enter the interactive mode according to the second control signal.
[0007] Furthermore, in the control mode, the step of acquiring the user's current image in real time through the image acquisition device and controlling the earphone to enter the corresponding dialogue mode based on the current image includes: When the headphones enter interactive mode, the image acquisition device captures the user's facial image in real time and performs image processing on the facial image to obtain the grayscale image corresponding to the facial image. Construct a coordinate system for the headphones, wherein the coordinate system includes a left coordinate system corresponding to the image acquisition device of the left headphone and a right coordinate system corresponding to the image acquisition device of the right headphone; Define the left imaging plane of the image acquisition device of the left earphone and the right imaging plane of the image acquisition device of the right earphone, and calculate all three-dimensional coordinates of the face contour in the face image based on the left coordinate system, the right coordinate system, the left imaging plane, the right imaging plane, and the grayscale image. The three-dimensional coordinates are verified using a preset database, and the headset is controlled to enter the corresponding dialogue mode based on the verification results.
[0008] Furthermore, the steps of acquiring the user's voice signal in real time through the voice acquisition device and preprocessing the voice signal to obtain the corresponding voice features include: The speech signal is pre-emphasized, and the pre-emphasized speech signal is then framed and windowed to obtain preliminary processed data. The pre-defined speech feature processing algorithm is used to extract speech features from the preliminary processed data to obtain the corresponding speech features.
[0009] Furthermore, the step of inputting the speech features into the speech emotion recognition model to identify the emotion type in the speech signal includes: The speech features are extracted using the deep convolutional neural network and long short-term memory network of the language emotion recognition model to obtain the corresponding deep frequency domain feature vector and deep temporal feature vector; The depth frequency domain feature vector and the depth temporal feature vector are fused to obtain the first fused feature; The speech features are processed using the parallel convolutional neural network in the language emotion recognition model to obtain speech temporal features and speech frequency features; The speech temporal features and speech frequency features are subjected to global average pooling, and the resulting processing results are adaptively fused in space to obtain a second fused feature. The first fusion feature and the second fusion feature are fused together, and the emotion type in the speech signal is identified based on the feature fusion result.
[0010] The present invention also proposes an artificial intelligence interaction system for headphones, wherein the headphones are equipped with at least an image acquisition device and a voice acquisition device, and the system includes: The status acquisition module is used to acquire the current status of the earphone and control the earphone to enter the corresponding control mode based on the current status; The mode control module is used to acquire the current image of the user of the headset in real time through the image acquisition device in the control mode, and control the headset to enter the corresponding dialogue mode according to the current image. The signal preprocessing module is used to collect the user's voice signal in real time through the voice acquisition device in the dialogue mode, and to preprocess the voice signal to obtain the corresponding voice features. An emotion recognition module is used to construct a language emotion recognition model and input the speech features into the speech emotion recognition model to identify the emotion type in the speech signal; The voice interaction module is used to obtain the corresponding voice template from a preset template database according to the emotion type, generate the corresponding interactive voice based on the voice template and the voice features, and realize the voice interaction between the headset and the user through the interactive voice.
[0011] Furthermore, the current state includes wearing state and idle state, the control mode includes noise cancellation mode, environment mode and interaction mode, and the state acquisition module includes: The state determination unit is used to control the headphones to enter the noise cancellation mode when the current state of the headphones is the wearing state. In the noise cancellation mode, the headphones will perform pulse suppression on the noise signals of the surrounding environment. The first mode switching unit is used to control the headphones to enter the ambient mode according to the first control signal when it receives the first control signal transmitted by the headphones in the noise cancellation mode. The second mode switching unit is used to control the headphones to enter the interactive mode according to the second control signal when it receives the second control signal transmitted by the headphones in the noise cancellation mode.
[0012] Furthermore, the mode control module includes: An image processing unit is used to acquire a user's facial image in real time through the image acquisition device when the earphone enters the interactive mode, and to perform image processing on the facial image to obtain a grayscale image corresponding to the facial image. A coordinate construction unit is used to construct the coordinate system of the earphone, wherein the coordinate system includes a left coordinate system corresponding to the image acquisition device of the left earphone and a right coordinate system corresponding to the image acquisition device of the right earphone; The coordinate calculation unit is used to define the left imaging plane of the image acquisition device of the left earphone and the right imaging plane of the image acquisition device of the right earphone, and to calculate all three-dimensional coordinates of the facial contour in the facial image based on the left coordinate system, the right coordinate system, the left imaging plane, the right imaging plane, and the grayscale image. The coordinate verification unit is used to verify the three-dimensional coordinates using a preset database, and control the earphone to enter the corresponding dialogue mode based on the verification result.
[0013] Furthermore, the signal preprocessing module includes: The signal processing unit is used to pre-emphasize the speech signal and perform frame segmentation and windowing on the pre-emphasized speech signal to obtain preliminary processed data. The first feature extraction unit is used to extract speech features from the preliminary processed data using a preset speech feature processing algorithm to obtain the corresponding speech features.
[0014] Furthermore, the emotion recognition module includes: The second feature extraction unit is used to extract features from the speech features using the deep convolutional neural network and long short-term memory network of the language emotion recognition model, so as to obtain the corresponding deep frequency domain feature vector and deep temporal feature vector. The first feature fusion unit is used to fuse the depth frequency domain feature vector with the depth temporal feature vector to obtain a first fused feature; The first feature processing unit is used to process the speech features using the parallel convolutional neural network in the language emotion recognition model to obtain speech time features and speech frequency features. The second feature fusion unit is used to perform global average pooling on the speech time features and the speech frequency features, and to adaptively fuse the processing results in space to obtain the second fused feature. The third feature fusion unit is used to fuse the first fusion feature and the second fusion feature, and identify the emotion type in the speech signal based on the feature fusion result.
[0015] The present invention also proposes a storage medium on which a computer program is stored, which, when executed by a processor, implements the aforementioned artificial intelligence interaction method for headphones.
[0016] The present invention also proposes an earphone, including a shell and an ear hook, wherein at least one image acquisition device and one voice acquisition device are provided on the ear hook, and a control module is provided inside the shell. When the control module executes a computer program, it implements the artificial intelligence interaction method of the earphone as described above.
[0017] The present invention discloses an AI interaction method, system, storage medium, and earphone for headphones. By collecting the current state of the earphone, the system controls the earphone to enter a corresponding control mode. In this control mode, it determines whether the user is a frequent user by collecting the user's current image. Preprocessing is performed on the voice signal from the earphone in dialogue mode to flatten the speech signal spectrum and improve the continuity of short-term signals. The voice features obtained from the voice signal are input into a corresponding voice emotion recognition model. Deep frequency domain feature vectors and deep temporal feature vectors are fused to achieve information interaction between features, improving information processing performance. The language emotion model is used to fuse speech time features and speech frequency features to capture complementary information between time and frequency features, reducing information redundancy and improving the model's recognition effect and accuracy. The identified emotion type is used to obtain a corresponding voice template. Interactive voice is generated using the voice features and voice template to achieve voice interaction between the earphone and the user. Attached Figure Description
[0018] Figure 1 This is a flowchart of the artificial intelligence interaction method for headphones according to the first embodiment of the present invention; Figure 2 This is a structural block diagram of the artificial intelligence interaction system of the headphones in the second embodiment of the present invention; Figure 3 This is a structural block diagram of the computer in the third embodiment of the present invention.
[0019] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation
[0020] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0022] Example 1 Please see Figure 1 The image shows an artificial intelligence interaction method for headphones according to the first embodiment of the present invention. The headphones are equipped with at least an image acquisition device and a voice acquisition device. The method specifically includes steps S101 to S105: S101, Obtain the current state of the earphone, and control the earphone to enter the corresponding control mode based on the current state; Furthermore, the current state includes wearing state and idle state, the control mode includes noise cancellation mode, environment mode and interaction mode, and step S101 specifically includes steps S1011~S1013: S1011, when the current state of the headphones is the wearing state, control the headphones to enter the noise cancellation mode, in which the headphones will pulse suppress the noise signal of the environment. S1012, when the first control signal transmitted by the headphones is received in the noise cancellation mode, the headphones are controlled to enter the ambient mode according to the first control signal; S1013, when the second control signal transmitted by the headphones is received in the noise cancellation mode, the headphones are controlled to enter the interactive mode according to the second control signal.
[0023] In this specific implementation, the image acquisition device is located on the earphone and facing the user's face. When the user wears the earphone, the sensor on the earphone triggers a corresponding signal to mark the current state of the earphone as the wearing state, and controls the earphone to enter the noise cancellation mode to achieve the effect of suppressing the ambient noise. Specifically, in the noise cancellation mode, the earphone will use a preset impulse noise suppression algorithm to suppress the noise in its environment, so as to improve the user's comfort when using the earphone and reduce the impact of ambient noise on the audio in the earphone.
[0024] Specifically, in noise cancellation mode, when the user triggers the first control signal, the headphones are controlled to enter the ambient mode. In this ambient mode, the noise suppression effect within a preset range (in this embodiment, the preset range is 1-2m) centered on the headphones is reduced, so that the user can collect ambient sounds within this range. The triggering method of the first control signal includes tapping the touch area on the earphone shell (e.g., tapping the touch area twice with an interval of no more than 1 second) and sending the first control signal directly to the headphones through the user's terminal device. Furthermore, in noise cancellation mode, when the user triggers the second control signal, the headphones enter the interactive mode. In this mode, the headphones will receive the audio signal emitted by the user in real time based on the noise cancellation mode, and generate corresponding interactive voice according to the user's audio signal and the preset interactive model, so as to realize the interactive effect between the user and the interactive model in the headphones. The triggering methods of the second control signal also include tapping the touch area on the outer shell of the headphones (e.g., tapping the touch area multiple times, with each tap interval not exceeding 1 second) and sending the second control signal directly to the headphones through the terminal device used by the user. S102, in the control mode, the image acquisition device acquires the current image of the user of the earphone in real time, and controls the earphone to enter the corresponding dialogue mode according to the current image; Furthermore, step S102 specifically includes steps S1021 to S1024: S1021, When the earphone enters the interactive mode, the image acquisition device acquires the user's facial image of the earphone in real time, and performs image processing on the facial image to obtain the grayscale image corresponding to the facial image. S1022, Construct the coordinate system of the earphone, wherein the coordinate system includes the left coordinate system corresponding to the image acquisition device of the left earphone and the right coordinate system corresponding to the image acquisition device of the right earphone; S1023, define the left imaging plane of the image acquisition device of the left earphone and the right imaging plane of the image acquisition device of the right earphone, and calculate all three-dimensional coordinates of the face contour in the face image based on the left coordinate system and the right coordinate system, the left imaging plane and the right imaging plane and the gray value image. S1024, the three-dimensional coordinates are verified using a preset database, and the headset is controlled to enter the corresponding dialogue mode based on the verification result.
[0025] In practice, when the headphones enter the interaction model, the current user's facial image is captured in real time by the image acquisition device, and the facial image is processed in grayscale to obtain the grayscale value image corresponding to the facial image. It can be understood that grayscale processing can better identify the contour information of the face in the image, so as to facilitate the subsequent filtering of the face coordinate information. Furthermore, a coordinate system for the headphones is constructed, comprising a left coordinate system corresponding to the image acquisition device of the left headphone and a right coordinate system corresponding to the image acquisition device of the right headphone. The left imaging plane of the left headphone's image acquisition device and the right imaging plane of the right headphone's image acquisition device are defined, with the origins of the left and right coordinate systems being respectively... , The left and right imaging planes are respectively , Define the spatial points of the final image (i.e., the edge points of the user's face marked using the grayscale image) as... This spatial point and the imaging plane , The intersection point is denoted as , Based on the left coordinate system, the right coordinate system, and the effective focal length of the image acquisition device , The corresponding imaging model is obtained: ; ; The positional relationship between the coordinate systems corresponding to the left and right image acquisition devices is transformed using a spatial geometric transformation matrix. express: ; The constructed spatial geometric transformation matrix is represented as follows: , where the rotation matrix Translation matrix The parameters of the rotation and translation matrices are obtained from the intrinsic parameter calibration of the image acquisition device. Spatial points are calculated based on the rotation and translation matrices and the aforementioned imaging model. 3D coordinates: ; Based on the above method, all three-dimensional coordinates of the face contour in the grayscale image are calculated. The entire three-dimensional coordinate system is then verified using a pre-set database to determine whether the current user is a frequent user. Understandably, this database stores the face contour coordinates of all users who have used the headphones. The calculated three-dimensional coordinates are compared with these face contour coordinates, and the user wearing the headphones is identified based on the comparison results. This allows the headphones to enter the corresponding dialogue mode for that user. In the dialogue mode, users can set the headphone functions to be activated, such as ambient sound acquisition, ambient noise reduction, and voice interaction.
[0026] S103, in the dialogue mode, the user's voice signal is collected in real time by the voice acquisition device, and the voice signal is preprocessed to obtain the corresponding voice features; Furthermore, step S103 specifically includes steps S1031 to S1032: S1031, the speech signal is pre-emphasized, and the pre-emphasized speech signal is framed and windowed to obtain preliminary processed data; S1032, the speech features of the preliminary processed data are extracted using a preset speech feature processing algorithm to obtain the corresponding speech features.
[0027] In practical implementation, in dialogue mode, the user's voice signal is collected in real time by a voice acquisition device. A high-pass filter is used to pre-emphasize the voice signal to flatten the spectrum of the voice signal and compensate for the energy attenuation caused by the high frequency band. The pre-emphasized voice signal is then subjected to Fourier transform, and framing and windowing processing are performed on the signal. Framing can divide the voice signal into short-time signals, and windowing can improve the continuity of short-time signals, thereby obtaining preliminary processed data.
[0028] Specifically, the preliminary data processing includes extracting speech features from the preliminary data using a preset speech feature processing algorithm (in this embodiment, the speech feature processing algorithm uses Fourier transform and Mel spectrum analysis algorithm). The preliminary data is then analyzed in the time domain and in the spectrum using the Fourier transform algorithm and the high-resolution processing algorithm, respectively. The obtained frequency domain features are then analyzed using the Mel spectrum analysis algorithm to obtain the speech features corresponding to the preliminary data.
[0029] In this embodiment, the aforementioned speech features include not only frequency domain features but also time domain features. The time domain features are achieved by performing high-resolution processing on the temporal information of the speech signal to segment the user's speech.
[0030] S104, Construct a language emotion recognition model and input the speech features into the speech emotion recognition model to identify the emotion type in the speech signal; Furthermore, step S104 specifically includes steps S1041 to S1045: S1041, the speech features are extracted using the deep convolutional neural network and long short-term memory network of the language emotion recognition model to obtain the corresponding deep frequency domain feature vector and deep temporal feature vector. S1042, the depth frequency domain feature vector and the depth temporal feature vector are fused to obtain the first fused feature; S1043, The parallel convolutional neural network in the language emotion recognition model is used to perform feature processing on the speech features to obtain speech time features and speech frequency features. S1044, global average pooling is performed on the speech time features and the speech frequency features, and the processing results are adaptively fused in space to obtain the second fused feature; S1045, perform feature fusion on the first fusion feature and the second fusion feature, and identify the emotion type in the speech signal based on the feature fusion result.
[0031] In specific implementation, a language emotion recognition model is constructed, which consists of multiple layers, including an input layer, a convolutional layer, a hidden layer, a max pooling layer, an average pooling layer, and an output layer. Each layer consists of multiple neurons. Specifically, the deep convolutional neural network in the language emotion recognition model is used to extract speech features to generate a deep frequency domain feature of size (256, 57). In the language emotion recognition model, the input layer copies the single-channel feature map into a three-channel tensor, and sets the size of the input layer to (3, 256, 57). The convolutional layer extracts features from the speech features, and the max pooling layer and average pooling layer compress the speech features. Each convolutional layer includes a first convolutional sub-layer and a second convolutional sub-layer. The kernel size of the first convolutional sub-layer is 5x5 and the number of kernels is 64. The kernel size of the second convolutional sub-layer is 3x3 and the number of kernels is 128. The average pooling layer compresses the depth frequency domain features of (256, 57) to obtain the corresponding depth frequency domain feature vector. Long Short-Term Memory (LSTM) networks of a language emotion recognition model are used to extract speech features. The LSM network has 256 hidden units, and the loss rate of the regularization layer used to prevent overfitting is 0.5. The regularization layer is fused with a fully connected layer, and the extracted features are input into the fused fully connected layer to obtain the corresponding deep temporal feature vector. The deep frequency domain feature vector and the deep temporal feature vector are adjusted to the same format, and the two vectors are weighted and fused to obtain the interaction information of the two features to obtain the first fused feature. The first fusion can distinguish the important feature regions and low-value feature regions in the emotion recognition process.
[0032] Furthermore, the speech features are processed using a parallel convolutional neural network in the language emotion recognition model. In the first network path of the parallel convolutional neural network, the speech features are processed by the first convolutional layer and output as a speech time feature. The size of the first convolutional layer is [16, 1x3]. In the second network path of the parallel convolutional neural network, the speech features are processed by the second convolutional layer and output as a speech frequency feature. The size of the first convolutional layer is [16, 1x3] and the size of the second convolutional layer is [16, 3x1]. Specifically, the speech time features and speech frequency features are weighted and transposed, and feature transformation is performed using the corresponding time and frequency dimensions respectively. Then, the global average pooling layer of the language emotion recognition model is used for feature compression to generate the corresponding speech time map and speech frequency map. The linear rectified function and hyperbolic tangent function are used as activation functions at the fully connected layer for weight learning, and multiplication is used to weight the speech time features and speech frequency features element by element to obtain new speech time features and speech frequency features. Specifically, the new speech time features and speech frequency features are concatenated in the channel direction, and the third convolutional layer (in this embodiment, the kernel size of the third convolutional layer is 1x1) is used to perform cross-channel interactive capture to obtain the corresponding weighted feature map. Furthermore, the Softmax algorithm is used to normalize the weight feature map and split it into corresponding spatial weights. The weight feature map is then weighted and fused according to the spatial weights. The speech time features and speech frequency features in the weight feature map are adaptively fused in space to obtain the second fused feature. Specifically, the first and second fusion features are fused to obtain the corresponding global features. The global features are then processed using a fully connected layer with a Softmax activation function to identify the final sentiment type.
[0033] S105, obtain the corresponding voice template from the preset template database according to the emotion type, generate the corresponding interactive voice based on the voice template and the voice features, and realize the voice interaction between the earphone and the user through the interactive voice.
[0034] In specific implementation, the corresponding voice template is obtained from the preset template database according to the emotion type, and the voice template and voice features are used to generate the corresponding interactive voice. Voice interaction between the headset and the user is realized by sending interactive voice. In this embodiment, the voice template includes the interactive voice template corresponding to the emotion type, including the fundamental frequency of the sound, the sound intensity, the sound harmonic noise ratio, and the sound pulse shape. This method can be pre-stored in the user's voice interaction database.
[0035] In summary, the AI interaction method for headphones in the above embodiments of the present invention collects the current state of the headphones, controls the headphones to enter the corresponding control mode based on the current state, and in the corresponding control mode, determines whether the user using the headphones is a frequent user by collecting the user's current image. Preprocessing is performed on the voice signal from the headphones in dialogue mode to flatten the spectrum of the voice signal and improve the continuity of short-term signals. The voice features obtained from the voice signal are input into the corresponding voice emotion recognition model, and the deep frequency domain feature vector and the deep temporal feature vector are fused to achieve information interaction between features and improve information processing performance. The language emotion model is used to fuse the voice time features and voice frequency features in the voice features to capture complementary information between the time features and frequency features, reduce information redundancy, and improve the recognition effect and accuracy of the model. The corresponding voice template is obtained using the recognized emotion type, and interactive voice is generated using the voice features and voice template to achieve voice interaction between the headphones and the user.
[0036] Example 2 In another aspect, this invention also proposes an artificial intelligence interaction system for headphones; please refer to [link / reference needed]. Figure 2 The image shows an artificial intelligence interaction system for headphones according to a second embodiment of the present invention. The headphones are equipped with at least an image acquisition device and a voice acquisition device. The system includes: The status acquisition module 11 is used to acquire the current status of the earphone and control the earphone to enter the corresponding control mode based on the current status; Furthermore, the current state includes wearing state and idle state, the control mode includes noise cancellation mode, environment mode and interaction mode, and the state acquisition module 11 includes: The state determination unit is used to control the headphones to enter the noise cancellation mode when the current state of the headphones is the wearing state. In the noise cancellation mode, the headphones will perform pulse suppression on the noise signals of the surrounding environment. The first mode switching unit is used to control the headphones to enter the ambient mode according to the first control signal when it receives the first control signal transmitted by the headphones in the noise cancellation mode. The second mode switching unit is used to control the headphones to enter the interactive mode according to the second control signal when it receives the second control signal transmitted by the headphones in the noise cancellation mode.
[0037] The mode control module 12 is used to acquire the current image of the user of the earphone in real time through the image acquisition device in the control mode, and control the earphone to enter the corresponding dialogue mode according to the current image. Furthermore, the mode control module 12 includes: An image processing unit is used to acquire a user's facial image in real time through the image acquisition device when the earphone enters the interactive mode, and to perform image processing on the facial image to obtain a grayscale image corresponding to the facial image. A coordinate construction unit is used to construct the coordinate system of the earphone, wherein the coordinate system includes a left coordinate system corresponding to the image acquisition device of the left earphone and a right coordinate system corresponding to the image acquisition device of the right earphone; The coordinate calculation unit is used to define the left imaging plane of the image acquisition device of the left earphone and the right imaging plane of the image acquisition device of the right earphone, and to calculate all three-dimensional coordinates of the facial contour in the facial image based on the left coordinate system, the right coordinate system, the left imaging plane, the right imaging plane, and the grayscale image. The coordinate verification unit is used to verify the three-dimensional coordinates using a preset database, and control the earphone to enter the corresponding dialogue mode based on the verification result.
[0038] The signal preprocessing module 13 is used to collect the user's voice signal in real time through the voice acquisition device in the dialogue mode, and to preprocess the voice signal to obtain the corresponding voice features. Furthermore, the signal preprocessing module 13 includes: The signal processing unit is used to pre-emphasize the speech signal and perform frame segmentation and windowing on the pre-emphasized speech signal to obtain preliminary processed data. The first feature extraction unit is used to extract speech features from the preliminary processed data using a preset speech feature processing algorithm to obtain the corresponding speech features.
[0039] The emotion recognition module 14 is used to construct a language emotion recognition model and input the speech features into the speech emotion recognition model to identify the emotion type in the speech signal; Furthermore, the emotion recognition module 14 includes: The second feature extraction unit is used to extract features from the speech features using the deep convolutional neural network and long short-term memory network of the language emotion recognition model, so as to obtain the corresponding deep frequency domain feature vector and deep temporal feature vector. The first feature fusion unit is used to fuse the depth frequency domain feature vector with the depth temporal feature vector to obtain a first fused feature; The first feature processing unit is used to perform feature processing on the speech features using the parallel convolutional neural network in the language emotion recognition model to obtain speech time features and speech frequency features. The second feature fusion unit is used to perform global average pooling on the speech time features and the speech frequency features, and to adaptively fuse the processing results in space to obtain the second fused feature. The third feature fusion unit is used to fuse the first fusion feature and the second fusion feature, and identify the emotion type in the speech signal based on the feature fusion result.
[0040] The voice interaction module 15 is used to obtain the corresponding voice template from a preset template database according to the emotion type, generate the corresponding interactive voice based on the voice template and the voice features, and realize the voice interaction between the headset and the user through the interactive voice.
[0041] The functions or operation steps implemented by the above modules and units are largely the same as those in the above method embodiments, and will not be repeated here.
[0042] The artificial intelligence interaction system for headphones provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the system embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0043] Example 3 This invention also proposes a computer, please refer to [link / reference]. Figure 3 The computer shown in the third embodiment of the present invention includes a memory 10, a processor 20, and a computer program 30 stored in the memory 10 and executable on the processor 20. When the processor 20 executes the computer program 30, it implements the above-described artificial intelligence interaction method for headphones.
[0044] The memory 10 includes at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 10 can be an internal storage unit of a computer, such as the computer's hard disk. In other embodiments, the memory 10 can be an external storage device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Furthermore, the memory 10 can include both internal and external storage units of the computer. The memory 10 can be used not only to store application software and various types of data installed on the computer, but also to temporarily store data that has been output or will be output.
[0045] In some embodiments, the processor 20 may be an electronic control unit (ECU), a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip, used to run program code stored in the memory 10 or process data, such as executing access restriction programs.
[0046] It should be pointed out that, Figure 3 The structure shown does not constitute a limitation on the computer. In other embodiments, the computer may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0047] This invention also proposes a storage medium storing a computer program that, when executed by a processor, implements the aforementioned artificial intelligence interaction method for headphones.
[0048] In this embodiment, the present invention also proposes an earphone, including a shell and an ear hook, wherein at least an image acquisition device and a voice acquisition device are provided on the ear hook, and a control module is provided inside the shell. When the control module executes a computer program, it implements the artificial intelligence interaction method of the earphone as described above.
[0049] The image acquisition device can be used to record videos of a certain specification. Users can interact with the screen of the housing (in this embodiment, the housing is a smart cabin with a touch screen), such as playing previews, wirelessly transmitting videos, etc. The image acquisition device can recognize people and the environment, and use the above-mentioned artificial intelligence interaction methods such as AI question answering, AI photography, and AI noise reduction. Furthermore, a G-Sensor is also installed on the shell, which can sense the user's wearing status and remind the user to stretch their neck when the user wears it for too long.
[0050] Those skilled in the art will understand that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0051] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0052] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0053] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0054] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. An artificial intelligence interaction method for headphones, wherein the headphones are equipped with at least an image acquisition device and a voice acquisition device, characterized in that, The method includes: Obtain the current state of the headphones, and control the headphones to enter the corresponding control mode based on the current state; In the control mode, the image acquisition device captures the user's current image in real time, and controls the earphone to enter the corresponding dialogue mode based on the current image; In the dialogue mode, the user's voice signal is collected in real time by the voice acquisition device, and the voice signal is preprocessed to obtain the corresponding voice features; A speech emotion recognition model is constructed, and the speech features are input into the speech emotion recognition model to identify the emotion type in the speech signal; According to the emotion type, the corresponding voice template is obtained from the preset template database, and the corresponding interactive voice is generated based on the voice template and the voice features. The interactive voice is used to realize the voice interaction between the headphones and the user. The current state includes both wearing state and idle state; the control mode includes noise cancellation mode, ambient mode, and interactive mode; and the step of obtaining the current state of the headphones and controlling the headphones to enter the corresponding control mode based on the current state includes: When the headphones are currently in a wearing state, the headphones are controlled to enter a noise cancellation mode. In the noise cancellation mode, the headphones will pulse suppress the noise signal of the surrounding environment. When a first control signal is received from the headphones in the noise cancellation mode, the headphones are controlled to enter the ambient mode according to the first control signal. When a second control signal is received from the headphones in the noise cancellation mode, the headphones are controlled to enter the interactive mode according to the second control signal.
2. The artificial intelligence interaction method for headphones according to claim 1, characterized in that, In the control mode, the steps of acquiring the user's current image in real time through the image acquisition device and controlling the earphone to enter the corresponding dialogue mode based on the current image include: When the headphones enter interactive mode, the image acquisition device captures the user's facial image in real time and performs image processing on the facial image to obtain the grayscale image corresponding to the facial image. Construct a coordinate system for the headphones, wherein the coordinate system includes a left coordinate system corresponding to the image acquisition device of the left headphone and a right coordinate system corresponding to the image acquisition device of the right headphone; Define the left imaging plane of the image acquisition device of the left earphone and the right imaging plane of the image acquisition device of the right earphone, and calculate all three-dimensional coordinates of the face contour in the face image based on the left coordinate system, the right coordinate system, the left imaging plane, the right imaging plane, and the grayscale image. The three-dimensional coordinates are verified using a preset database, and the headset is controlled to enter the corresponding dialogue mode based on the verification results.
3. The artificial intelligence interaction method for headphones according to claim 1, characterized in that, The steps of acquiring the user's voice signal in real time through the voice acquisition device and preprocessing the voice signal to obtain the corresponding voice features include: The speech signal is pre-emphasized, and the pre-emphasized speech signal is then framed and windowed to obtain preliminary processed data. The pre-defined speech feature processing algorithm is used to extract speech features from the preliminary processed data to obtain the corresponding speech features.
4. The artificial intelligence interaction method for headphones according to claim 1, characterized in that, The steps of inputting the speech features into the speech emotion recognition model to identify the emotion type in the speech signal include: The speech features are extracted using the deep convolutional neural network and long short-term memory network of the speech emotion recognition model to obtain the corresponding deep frequency domain feature vector and deep temporal feature vector. The depth frequency domain feature vector and the depth temporal feature vector are fused to obtain the first fused feature; The speech features are processed using the parallel convolutional neural network in the speech emotion recognition model to obtain speech time features and speech frequency features. The speech temporal features and speech frequency features are subjected to global average pooling, and the resulting processing results are adaptively fused in space to obtain a second fused feature. The first fusion feature and the second fusion feature are fused together, and the emotion type in the speech signal is identified based on the feature fusion result.
5. An artificial intelligence interaction system for headphones, wherein the headphones are equipped with at least an image acquisition device and a voice acquisition device, characterized in that, The system includes: The status acquisition module is used to acquire the current status of the earphone and control the earphone to enter the corresponding control mode based on the current status; The mode control module is used to acquire the current image of the user of the headset in real time through the image acquisition device in the control mode, and control the headset to enter the corresponding dialogue mode according to the current image. The signal preprocessing module is used to collect the user's voice signal in real time through the voice acquisition device in the dialogue mode, and to preprocess the voice signal to obtain the corresponding voice features. An emotion recognition module is used to construct a speech emotion recognition model and input the speech features into the speech emotion recognition model to identify the emotion type in the speech signal; The voice interaction module is used to obtain the corresponding voice template from a preset template database according to the emotion type, generate the corresponding interactive voice based on the voice template and the voice features, and realize the voice interaction between the headset and the user through the interactive voice. The current state includes wearing state and idle state; the control mode includes noise cancellation mode, environment mode, and interaction mode; and the state acquisition module includes: The state determination unit is used to control the headphones to enter the noise cancellation mode when the current state of the headphones is the wearing state. In the noise cancellation mode, the headphones will perform pulse suppression on the noise signals of the surrounding environment. The first mode switching unit is used to control the headphones to enter the ambient mode according to the first control signal when it receives the first control signal transmitted by the headphones in the noise cancellation mode. The second mode switching unit is used to control the headphones to enter the interactive mode according to the second control signal when it receives the second control signal transmitted by the headphones in the noise cancellation mode.
6. The artificial intelligence interaction system for headphones according to claim 5, characterized in that, The mode control module includes: An image processing unit is used to acquire a user's facial image in real time through the image acquisition device when the earphone enters the interactive mode, and to perform image processing on the facial image to obtain a grayscale image corresponding to the facial image. A coordinate construction unit is used to construct the coordinate system of the earphone, wherein the coordinate system includes a left coordinate system corresponding to the image acquisition device of the left earphone and a right coordinate system corresponding to the image acquisition device of the right earphone; The coordinate calculation unit is used to define the left imaging plane of the image acquisition device of the left earphone and the right imaging plane of the image acquisition device of the right earphone, and to calculate all three-dimensional coordinates of the facial contour in the facial image based on the left coordinate system, the right coordinate system, the left imaging plane, the right imaging plane, and the grayscale image. The coordinate verification unit is used to verify the three-dimensional coordinates using a preset database, and control the earphone to enter the corresponding dialogue mode based on the verification result.
7. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the artificial intelligence interaction method for the headphones as described in any one of claims 1 to 4.
8. An earphone, comprising a housing and an ear hook, characterized in that, At least one image acquisition device and one voice acquisition device are provided on the ear hook, and a control module is provided inside the housing. When the control module executes a computer program, it implements the artificial intelligence interaction method of the headphones as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Smart headphone
CN107404682A
Scene and emotion oriented Chinese speech synthesizing method and device, and storage medium
CN110211563A
Earphone control method, earphone and computer readable storage medium
CN113873382A