An interactive earphone and system thereof
By using a generative adversarial network that fuses bone conduction signals and lip motion images, combined with visual and physiological information processing, the robustness of multimodal fusion and intelligent interaction of headphone systems in extreme environments are solved, achieving efficient speech recognition and multimedia interaction to meet the diverse needs of special scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2026-03-27
AI Technical Summary
Existing headphone systems lack robustness in multimodal fusion under extremely complex environments, have limited applicability, and low levels of intelligent interaction, failing to meet diverse needs and tasks while ensuring portability and wearability.
Design an interactive headset system that combines a perception module, a computing module, and a display module. Employ a method of fusing bone conduction signals and lip motion images, and achieve deep fusion of multimodal information through generative adversarial networks and cross-modal attention mechanisms. Integrate visual and physiological information processing to support 3D immersive video calls and remote physical and mental state monitoring.
It achieves efficient voice recognition and multimedia interaction in complex environments, improves the applicability and intelligent interaction of the headphone system, meets the diverse needs of special scenarios, and provides portability and real-time monitoring functions.
Smart Images

Figure CN116095548B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of human-computer interaction and speech recognition, and particularly relates to an interactive earphone and a system thereof. BACKGROUND
[0002] As a remote information interaction tool with audio as the core, earphones have mature baseline technology and extensive application needs. In daily applications, benefiting from the stable scene, controllable noise, and small environmental factor interference, etc., there are already system designs that can maturely realize basic functions. In special scenarios such as military operations, emergency rescue, field exploration, medical rehabilitation, etc., or in response to special needs such as outdoor hiking, near-sea navigation, self-driving travel, etc., due to reasons such as high noise interference, strong concealment needs, and the health status of the sound producer, the traditional air-guided audio channel cannot fully represent the speech information, and even completely loses the representation ability. Therefore, on the basis of earphone systems based on daily applications, on the one hand, anti-noise and sound pickup design is needed, and on the other hand, multi-modal information such as lip language, muscle electrical signals, and bone conduction audio is needed to be integrated to cooperatively model, so as to ensure the use effect of the earphone in special scenarios. There are currently related earphone systems that can meet the basic communication needs, have certain active noise reduction design and intelligent directional sound pickup, and also have multi-modal fusion-based multi-mode solutions, i.e., users can select different modal combinations to pick up speech information according to the actual situation.
[0003] However, whether for special application scenarios or for daily special needs, the current earphone systems still have two deficiencies: one is that the multi-modal fusion method is simple, mainly showing low robustness to external disturbances in extremely complex environments, small range of applicable speech models, and low application extensibility; the other is that the intelligent interaction degree is insufficient, mainly showing single function, low integration, low module utilization, and inability to balance the use of portability and wearability while meeting multi-element needs and tasks, etc.
[0004] The first point is that the multi-modal fusion method is simple, mainly showing low robustness to external disturbances in extreme complex environments, small applicable voice model range, and low application extensibility.
[0005] The second point is that the intelligent interaction degree is insufficient, mainly showing single function, low integration, and low module utilization rate, which cannot meet multiple demands and tasks while ensuring portability and wearing friendliness, specifically, in actual communication demands, audio is the most core information carrier, but introducing more-dimensional information interconnection can provide better experience for both parties or multiple parties in communication. Such multi-dimensional information, such as video calls, enables each terminal to see the face pictures and scenes of each other, which is an upgrade of interaction degree from the visual information dimension. For example, by means of sensing technology, each terminal can see and understand the physical and mental signs of each other, which is an upgrade of interaction degree from the physiological information dimension. At the same time, the introduction of similar multi-dimensional information not only provides users with better interactive experience, but also is a necessary function in application scenarios such as special operations, rescue, exploration, and can help the main control terminal to understand the health status of individuals in real time in group communication, so as to make corresponding judgments and decisions. The existing earphone system only focuses on basic functions, that is, the interaction of audio information, and cannot realize the function of multimedia interaction as described above. If the earphone system is separated, and a camera system or a multi-sensing system is configured to undertake the above function, the wearing friendliness and use portability in mobile application scenarios will be greatly reduced. Therefore, in the earphone system taking audio information as the leading interaction mode, designing an intelligent system integrated with multimedia interaction function to provide a solution with multiple functions and portability is lacking in the current earphone system design. SUMMARY
[0006] Based on the defects of the prior art, the present application aims to design an interactive earphone and system with high integration and good portability. The earphone can be applied to audio-visual communication and health monitoring of patients with voice disorders in medical and rehabilitation engineering, remote communication and live monitoring in high-noise or covert private military operations and emergency rescue scenarios, and can also be used for other daily audio-visual communication.
[0007] To achieve the above technical purposes, the technical scheme adopted by the present application is as follows:
[0008] An interactive earphone system comprises a perception module, a calculation module, a communication module and a display module.
[0009] The perception module is used to acquire information of various modalities; the perception module comprises an audio information perception unit and a visual information perception unit.
[0010] The calculation module is used to process the information of various modalities to obtain a processing result; the calculation module comprises an audio information processing unit, a visual information processing unit and a physiological information processing unit.
[0011] The communication module is used to transmit the processing result of the calculation module to other terminals; receive information from other terminals and send it to the display module.
[0012] The display module is used to present auditory information and visual information; the display module comprises a loudspeaker and an augmented reality glasses.
[0013] Further, the audio information perception unit of the perception module comprises a microphone and a bone conduction signal sensor.
[0014] The visual information perception unit comprises a first visual perception unit, a second visual perception unit and a third visual perception unit; the first visual perception unit acquires a user's facial image, the second visual perception unit acquires a scene image, and the third visual perception unit acquires a user's eye movement image.
[0015] Further, the air conduction audio signal is acquired by the microphone, the lip image signal is acquired by the first visual information perception unit, and the bone conduction signal is acquired by the bone conduction signal sensor; the acquired signals are calculated by the audio information processing unit, recognized as specific phrases and instructions, and simultaneously synthesized into clear audio, realizing audio interaction.
[0016] The user's facial image is acquired by the first visual perception unit, and the scene image is acquired by the second visual perception unit; the acquired facial image and scene image are processed by the visual information processing unit, realizing regular video call or three-dimensional immersive video call.
[0017] The user's facial image is acquired by the first visual perception unit, and the eye movement image is acquired by the third visual perception unit; the acquired facial image and eye movement image are calculated by the physiological information processing unit, obtaining the user's heart rate and recognizing the user's emotional type, realizing monitoring of the user's physical and mental state.
[0018] Further, the audio information processing unit directly takes bone conduction audio as the target modality for sound pickup in a quiet and stable environment, converts and transmits it through a traditional earphone system, and realizes audio interaction.
[0019] In a complex scene with high noise and high mobility, the audio information processing unit takes the lip image signal and the bone conduction signal as the target modality for sound pickup, and realizes modal information fusion audio interaction by the bone conduction-lip reading fusion-synthesis method.
[0020] The bone conduction-lip reading fusion-synthesis method is as follows:
[0021] The bone conduction voice signal and the lip movement image signal synchronously acquired when collecting the user voice input are acquired.
[0022] The single-modality data features in the time domain and the spatial domain are determined based on the bone conduction voice signal and the lip movement image signal.
[0023] Based on the determined single-modality data features in the time domain and the spatial domain, a generative adversarial network with cross-modal attention mechanism and a mel-spectrogram fusion method are applied to obtain modal collaborative feature expression.
[0024] Based on the obtained modal collaborative feature expression, on the one hand, a trained back-end classification neural network model is applied to output specific phrases and instructions; on the other hand, a vocal synthesis model is applied to obtain an audio waveform.
[0025] Further, the visual information processing unit includes a regular computing unit and a three-dimensional immersive computing unit, and the visual information processing unit is divided into a regular mode and a three-dimensional immersive mode.
[0026] In the regular mode, a regular two-dimensional image-based video call is realized; the regular computing unit corrects the distortion of the face image, and the ideal pixel point coordinates (x, y) and the distorted pixel point coordinates (x d , y d ) have the following relationship
[0027] x d =x+x[k1(x 2 +y 2 )+k2(x 2 +y 2 ) 2 ]
[0028] y d =y+y[k1(x 2 +y 2 )+k2(x 2 +y 2 ) 2 ]
[0029] Wherein k1, k2 is radial distortion coefficient, using Zhang Zhengyou calibration method to calculate the radial distortion coefficient, according to the relationship of the above formula, the ideal pixel point (x, y) can be solved by anti-distortion calculation from the distorted pixel point, that is, the corrected image can be obtained, and the video call based on two-dimensional image is realized;
[0030] In three-dimensional immersion mode, the calculation unit in the conventional mode still works, the distortion of the image is corrected first, and then the corrected image is processed by the three-dimensional immersion calculation unit, so that the real-time three-dimensional reconstruction of the portrait is realized, and three-dimensional immersion video call is realized.
[0031] Further, the three-dimensional immersion mode calculation comprises:
[0032] The face of the user is scanned by a structured light system, and by adjusting the illumination, the three-dimensional contour features of the user in a static and expressionless state under a certain illumination intensity interval and illumination angle interval are obtained, and (u, v) represents different intensity intervals and illumination angles.
[0033] V (u,v) ={X1, Y1, Z1, X2, Y2, Z2,..., X n , Y n , Z n}∈R 3n
[0034] The mapping relationship Z i =αX i +βY i , i = 1, 2,..., n
[0035] At the same time, the two-dimensional facial feature point data V j of the user under different basic expression categories (including anger, disgust, fear, sadness, expectation, joy, surprise, trust, etc.) is obtained, and j is the expression category index.
[0036] V j ={x1, y1, x2, y2,..., x n , y n}∈R 2n
[0037] The extended three-dimensional facial feature data V j-e can be obtained.
[0038] z i =αx i +βy i , i = 1, 2,..., n
[0039] V j- e = {x1, y1, z1, x2, y2, z2,..., x n , yn , z n}∈R 3n
[0040] The static expressionless shape feature is combined with the extended 3D face feature data randomly to form a shape basis vector S, as follows:
[0041] r (j-e) ~ Bernoulli(p)
[0042] V j-e ' = r (j-e) * V j-e
[0043] S (u,v) = V (u,u) + W j-e V j-e
[0044] (where the wavy line above the symbol has no special meaning, and is only for distinguishing from the foregoing);
[0045] r is n 0 or 1 randomly generated according to a Bernoulli distribution with a probability p. Then multiplied by V j-e相 , to randomly discard or retain a part of the extended 3D face feature vector, and multiplied by the weight of each extended 3D face feature vector, and then added to the static expressionless shape feature; W j-e is the weight of each group of 3D vectors.
[0046] Meanwhile, the face of the user is scanned in 3D, and the static expressionless texture feature T (u,u) of the user is obtained,
[0047] T (u,v) = {R1, G1, B1, R2, G2, B2, …, R n , G n , B n}∈R 3n
[0048] R, G, and B represent color components;
[0049] The above data matrices S (u,v) and T (u,v) are reduced in dimension through principal component analysis, and two principal component analysis models of the shape feature and the texture feature are obtained, respectively:
[0050]
[0051]
[0052] The mean values of the two shape features and the texture features respectively,
[0053] V S = [v s1 , v s2 , v s3 , …, v sm ] ∈ R 3n*m , V T = [v t1 , v t2 , v t3 , …, v tm ] ∈ R 3n*m , V S , V T are m principal components of the shape feature S (u,v) and the texture feature T (u,v) respectively, σ ∈ R m represents the standard deviation; accordingly, the three-dimensional face model including the shape model and the color model is obtained as follows, and the two models are superimposed to obtain the final modeling result I;
[0054] Shape model:
[0055] Texture model:
[0056] Three-dimensional reconstruction model: I = S + T
[0057] where λ i and ρ i represent the shape parameter and the texture parameter respectively;
[0058] The shape and texture principal component analysis models M S , M T of the obtained base vector group are stored after being labeled according to the scanned user, and a face base model database is obtained;
[0059] In the three-dimensional immersion mode, the key feature parameters of the two-dimensional face image obtained by the visual perception unit are identified in real time, the selection of the group of feature parameters is consistent with the aforementioned established two-dimensional face feature point data V j of the user, and the corresponding shape parameter λ i is obtained based on the aforementioned prior model;
[0060] The key feature points affected by the lighting condition are selected again, and the texture parameter ρ i constrained by the lighting condition is solved by using the bilinear interpolation method;
[0061] Based on the real-time obtained face base vector group, and shape parameters and texture parameters, the two-dimensional face image is mapped into a three-dimensional face image in real time by using the three-dimensional reconstruction model;
[0062] Further, the physiological information processing unit processes the face image obtained by the first visual perception unit, measures the heart rate of the user by using a remote photoplethysmography method, and transmits the heart rate measurement result to the display module.
[0063] Further, the physiological information processing unit processes the face image obtained by the first visual perception unit and the eye movement information obtained by the second visual perception unit.
[0064] For the face image, first, face alignment and normalization preprocessing are performed, face image features and time sequence information are extracted through a convolutional neural network (CNN) and a long short-term memory (LSTM), and then deep features are obtained, and a multi-classification result R0 is output through a shallow classifier; the face image features include face feature points, space-time domain features of facial micro-expressions, and optical flow features, etc.
[0065] For the eye movement information, eye movement features are obtained, including pupil diameter, gaze deviation, gaze duration, saccade duration, saccade amplitude, blink duration, and blink frequency; wherein the principal component analysis is performed on the pupil diameter information, the feature information is smoothed and normalized, the automatic encoder based on the restricted Boltzmann machine is used to encode the eye movement features according to the weight, the high-order feature expression is extracted, and the multi-classification result R1 is output through the shallow classifier.
[0066] The multi-classification result is the recognized different emotion categories, such as anger, disgust, fear, sadness, anticipation, joy, surprise, and trust, etc.
[0067] The preliminary emotion recognition results R0 and R1 are obtained by the above method; meanwhile, the features of the two channels are fused in the feature layer, i.e., the face image features and the eye movement features are normalized and spliced, and then the emotion recognition result R2 is output through the back-end classifier; the preliminary emotion recognition results R0 and R1 and the fusion emotion recognition result R2 are fused in the decision layer to obtain the final reliable emotion recognition result R.
[0068] The application also provides an interactive earphone, which comprises the interactive earphone system; the interactive earphone system comprises a microphone, a bone conduction sensor, a face camera, a loudspeaker, and an augmented reality glasses.
[0069] The outer side of the augmented reality glasses is provided with a scene camera, and the inner side of the augmented reality glasses is provided with a macro camera.
[0070] Further, the face camera is used as a first visual information sensing unit to capture a face image of the user in real time, the scene camera is used as a second visual information sensing unit to capture a scene image in real time, and the micro-distance camera is used as a third visual information sensing unit to capture eye movement information of the user in real time.
[0071] Compared with the prior art, the present application has the following beneficial effects:
[0072] On the one hand, the present application solves the technical problems of the prior art, such as simple multi-modal fusion method, low applicability in complex environment, poor generalization, and low application extension, by grafting an optical capture module into a bone conduction-lip reading fusion-synthesis system architecture to establish a speech synthesis solution under a modal information deep fusion mechanism.
[0073] On the other hand, the optical capture module is reused and integrated, and an augmented reality glasses is used at the same time to realize functions such as regular video call or three-dimensional immersive video call, remote physical and mental state monitoring, and the like, so as to expand the functions of the earphone system, improve the interactive experience of the user, and meet the needs of real-time monitoring of the real-time status and physical and mental indications of the user in special scenarios, thereby solving the technical problems of the prior art, such as single function and insufficient intelligent interaction, which cannot meet the needs of the special scenarios.
[0074] The interactive earphone system of the present application realizes special speech recognition, audio and video call, and physical and mental state monitoring in daily use or in high-noise and high-mobility scenarios while ensuring the portability of the headset. Based on a single headset, the system architecture and software and hardware design are highly integrated to realize intelligent remote audio interaction, remote video interaction, remote physiological information interaction, and other multimedia fusion functions. BRIEF DESCRIPTION OF DRAWINGS
[0075] Figure 1 Fig. 1 is a schematic diagram of the overall framework of the interactive earphone system;
[0076] Figure 2 Fig. 4 is an example diagram of the content presented by the display module;
[0077] Figure 3 Fig. 5 is a flowchart of the processing of audio information;
[0078] Figure 4 Fig. 6 is a schematic diagram of the division of the Mel spectrogram region;
[0079] Figure 5 Fig. 7 is a schematic diagram of the secondary partitioning of the low-frequency region in the Mel spectrogram;
[0080] Figure 6 Fig. 8 is a flowchart of the speech synthesis method based on the fusion of bone conduction signals and lip images;
[0081] Figure 7 A processing flowchart for visual information;
[0082] Figure 8 A processing flowchart for physiological information. DETAILED DESCRIPTION
[0083] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the embodiments of the present application will be specifically introduced below in combination with the drawings. Obviously, the following description only relates to some embodiments of the present application, and is not a limitation on the present application.
[0084] An interactive earphone system, comprising a perception module, a calculation module, a communication module and a display module, the overall framework is as shown in Figure 1
[0085] The perception module: multi-dimensional perception of the information released by the user and the on-site environment information, the perception module includes an audio information perception unit and a visual information perception unit, and the perception module transmits the information of each mode obtained to the calculation module;
[0086] In the field of computer human-computer interaction, mode refers to the existence form of data, such as different file formats of text, audio, image, video, etc. The modes involved in the present application mainly include audio, image and video.
[0087] The audio information perception unit of the perception module includes a microphone and a bone conduction signal sensor;
[0088] The visual information perception unit of the perception module includes a first visual perception unit, a second visual perception unit and a third visual perception unit; the first visual perception unit acquires a user facial image, the second visual perception unit acquires a scene image, and the third visual perception unit acquires a user eye movement image.
[0089] The calculation module: accepting the information transmitted by the perception module, performing calculation and processing on the information data, giving corresponding processing results, and transmitting the processing results to the display module and the communication module;
[0090] The calculation module includes an audio information processing unit, a visual information processing unit and a physiological information processing unit;
[0091] The calculation module further includes a main control unit for controlling the whole system and a storage unit for storing information;
[0092] The calculation module is the core module for realizing the functions of the system, adopts a system-level chip with a high-computing-power embedded neural network processor, deploys a neural network model, sets the above-mentioned processing units, and performs identification and calculation on the audio information, visual information and physiological information respectively.
[0093] Specifically, the air conduction audio signal acquired by the microphone, the lip image signal acquired by the first visual information sensing unit, and the bone conduction signal acquired by the bone conduction signal sensor are calculated by the audio information processing unit, recognized as specific phrases and instructions, and synthesized into clear audio at the same time, realizing audio interaction.
[0094] Specifically, the face image acquired by the first visual sensing unit and the scene image acquired by the second visual sensing unit are calculated by the visual information processing unit to realize video call.
[0095] Specifically, the face image acquired by the first visual sensing unit and the eye movement image acquired by the third visual sensing unit are calculated by the physiological information processing unit to obtain the heart rate of the user and identify the emotional type of the user, realizing the monitoring of the physical and mental state of the user.
[0096] Communication module: receiving the calculation results of the operation module, transmitting to other terminals; and receiving information from other terminals and sending to the display module;
[0097] Display module: taking the loudspeaker and augmented reality glasses as the display module, respectively presenting the auditory information and visual information.
[0098] The audio of other terminals is transmitted to the display module of the terminal through the communication module and played through the loudspeaker. The augmented reality glasses receive the information from other terminals transmitted by the communication module through Bluetooth or wireless network communication, and at the same time receive the calculation results transmitted by the physiological information processing unit in the operation module of the terminal.
[0099] Optionally, in the video call, the information displayed by the display module mainly includes the face video of the interlocutor or the environment video, and at the same time, the physical and mental indicators calculated by the physiological information processing unit of the opposite terminal (including the current emotional state and heart rate of the interlocutor), and the physical and mental indicators calculated by the physiological information processing unit of the terminal (including the current emotional state and heart rate of the user) are presented in the picture.
[0100] Optionally, the face video of the interlocutor and the environment video are respectively the main picture and the small window picture visible to the user by default, and the user can switch the main picture and the small window picture through voice instructions or eye gaze.
[0101] Optionally, the physical and mental indicators of the interlocutor are displayed by default in a specific position of the main picture, and the user can switch them to the physical and mental indicators of the user through voice instructions or eye gaze.
[0102] An example of the information presented by the display module is shown in Figure 2 .
[0103] The processing flow of the audio information is shown in Figure 3The specific implementation method of the audio interaction function is as follows:
[0104] In a quiet and stable environment, since the air conduction audio is not disturbed greatly, the audio information processing unit can directly take the air conduction audio as a target modality for sound pickup, and after conversion and transmission through a traditional earphone system, the function of audio interaction can be realized.
[0105] In a complex scene with high noise and high mobility, the air conduction channel will be blocked, so the audio information processing unit takes the lip image signal and the bone conduction signal as the target modality for sound pickup, and uses the bone conduction-lip reading fusion-synthesis method to realize modal information fusion speech recognition and speech synthesis.
[0106] The bone conduction-lip reading fusion-synthesis method includes the following steps:
[0107] S1, synchronously acquiring the bone conduction speech signal and the lip movement image signal when collecting user voice input;
[0108] S2, determining single-modality data features in the time domain and the spatial domain based on the bone conduction speech signal and the lip movement image signal;
[0109] S3, based on the determined single-modality data features in the time domain and the spatial domain, applying a generative adversarial network with cross-modal attention mechanism and a mel-spectrogram fusion method to establish a speech model and obtain modal collaborative feature expression;
[0110] Applying a time-frequency partition image fusion method of mel-spectrogram to establish a speech model and obtain modal collaborative feature expression of common information;
[0111] S4, based on the obtained modal collaborative feature expression, on the one hand, applying a trained back-end classification neural network model to identify and output specific phrases and instructions; on the other hand, applying a vocal synthesis model to obtain an audio waveform.
[0112] The following describes each link of the above steps S2, S3, and S4 in detail:
[0113] Step S2 specifically includes:
[0114] Step S21, after obtaining the bone conduction signal, processing (feature extraction) is performed to obtain a mel-spectrogram Mel-BC based on the bone conduction signal;
[0115] The bone conduction signal adopts an existing processing technique, that is, pre-emphasis, windowing, framing, short-time Fourier transform (STFT), obtaining a power spectrum, applying spectral subtraction for noise reduction processing, using a mel filter bank, obtaining a mel spectrogram, and extracting mel frequency cepstrum coefficients (MFCC); the obtained mel spectrogram Mel-BC will be one of the input channels in the subsequent spectrogram-based synthesis;
[0116] Step S22, obtain a sequence of frame pictures of the lip movement image, perform motion detection, input the sequence of frame picture data streams to the front-end neural network model, and extract the lip image feature F v The front-end neural network model can specifically use a VGG19 or Resnet-50 convolutional neural network model.
[0117] Thus, two single-modal data features are obtained.
[0118] Step S3 specifically includes:
[0119] Step S31, blind enhancement is performed on the bone conduction signal, and a certain high-frequency component is restored by using a traditional signal processing technique; specifically:
[0120] First, the blind enhancement based on spectral envelope conversion is performed on the bone conduction signal, and the high-frequency component is preliminarily filled (that is, preliminarily expanded) in a coarse-grained manner. Blind enhancement means that only the remaining low-frequency signal is used to restore the high-frequency component. The blind enhancement based on spectral envelope conversion is to use a pre-trained model to map the spectral envelope feature of the bone conduction signal to the spectral envelope feature of the air conduction signal, and then perform convolution calculation to obtain a complete signal with the estimated high-frequency component expanded. Based on this signal, the above-mentioned method of extracting the bone conduction signal data feature is repeated to obtain the corresponding mel frequency cepstrum coefficient as the feature representation Fb of the bone conduction signal in this link;
[0121] Step S32, the blind-enhanced bone conduction signal and the lip movement image signal are subjected to collaborative representation based on a cross-modal attention mechanism, and are input into a trained generative adversarial network to generate a mel spectrogram Mel-Vba based on modal fusion after multiple iterations; specifically:
[0122] First, around the commonly used Chinese short sentence, text instructions, logical instructions, based on the generative adversarial network (Generative Adversarial Nets), train the speech model, and construct the mapping relationship between the short sentence command and the corresponding mel spectrum. The training of the generative adversarial network can be regarded as the training of the generator (Generator) and the classifier (Discriminator) respectively. Among them, the generator generates a simulation picture based on the to-be-processed features, and in the design, the to-be-processed features refer to the cooperative features based on the lip movement image and the bone conduction signal, and a real and reliable mel spectrum is generated to supply the subsequent mel spectrum of the bone conduction signal channel Mel-BC and synthesize the speech accordingly. The classifier can give a consistency evaluation with the prior real picture according to the simulation picture output by the generator. For example, the mel spectrum generated based on the cooperative features is judged by the classifier to obtain two types of results: one is the unconditional result, that is, whether the mel spectrum is real or not, and the other is the conditional result, such as the similarity K1 between the mel spectrum and the prior real mel spectrum of sentence A, the similarity K2 between the mel spectrum and the prior real mel spectrum of sentence B, and the similarity K3 between the mel spectrum and the prior real mel spectrum of sentence C.
[0123] At the same time, in order to obtain a generated mel spectrum with high authenticity and high accuracy, the scheme requires that the features input into the generative adversarial network (that is, the input of the generator) contain as much effective information of the two modalities as possible. Therefore, the cross-modal attention mechanism is introduced in this embodiment to cooperatively express the input features, and the effective information expression is selected between the bone conduction signal feature and the lip movement image signal feature according to the weight, and the secondary information is eliminated. Modality refers to the form of data, such as text, audio, image, video, and different file formats are different modalities. The cross-modal task can study the association and connection between different modalities, so as to well integrate and process the information of the two modalities.
[0124] The cooperative expression method using the cross-modal attention mechanism is as follows:
[0125] The query vector, the key vector, and the value vector are respectively. The bone conduction signal original feature is calculated by a flattening operator of the frequency dimension and the channel dimension of the combined speech representation, and multiplied by the query vector to obtain The lip image original feature is multiplied by the key vector and the value vector respectively to obtain Thereafter, the weight is calculated according to the following formula;
[0126]
[0127]
[0128]
[0129] The obtained F is combined with the lip image original feature F v Concat splicing is performed, and the required feature expression is obtained.
[0130] On the other hand, it cannot be considered that the mel-spectrogram obtained by the generator only once can have satisfactory performance on both unconditional results and conditional results, and the authenticity and accuracy always have room for improvement. Therefore, the embodiment introduces an iterative mechanism, sets multiple orders of generators and classifiers, guides the iteration through the evaluation results of the classifier, and constantly updates the input features of each order generator through the cross-modal attention mechanism, gradually improves the feature purity, thereby improving the authenticity of the generated mel-spectrogram, and ensuring the accurate mapping relationship, so that the subsequent mel-spectrogram to speech waveform conversion is more realizable.
[0131] The following will specifically describe the visual-bone conducted attentional GAN (Visual-Bone Conducted Attentional GAN) model Vba-GAN designed based on the bone conduction-lip reading attention mechanism:
[0132] Step S321: The feature representation extracted based on the lip movement image signal is denoted as F v , as the original input I0;
[0133] Step S322: The original input I0 and the feature representation F b extracted based on the bone conduction signal are subjected to collaborative coding based on the cross-modal attention mechanism (Visual-Bone Conducted Attention);
[0134] Step S323: The weight feature F a1 after the collaborative coding based on the cross-modal attention mechanism is spliced with the original input I0 to form a fusion feature F c1 , which is input into the current order generator GE1 to obtain a generated feature F ml , and a mel-spectrogram IM1 of the current order is generated at the same time;
[0135] Step S324: The mel-spectrogram of the current order is input into the current order classifier D1, if the unconditional result of the judgment is true value and the conditional result of the judgment exists similarity K h of a certain sentence higher than a certain threshold (set to 90% in the embodiment), and the highest similarity K s of the remaining sentences is lower than a certain threshold (set to 3% in the embodiment), it is considered that the mel-spectrogram generated by the current order is a usable mel-spectrogram;
[0136] Step S325: If the judgment condition in step S324 is not met, the generated current order feature is expressed as F m1 As the original input I0 in step S321, according to steps S322, S323, S324, S325, the loop iteration is performed, and F a1 , F c1 , F m1 , GE1, IM1, D1 are incremented by 1;
[0137] Step S326: Until the judgment condition in S234 is met, end the iteration, output the generated available mel-spectrogram of the current order, thereby obtaining the mel-spectrogram Mel-Vba based on the collaborative coding of the generative adversarial network and the cross-modal attention mechanism.
[0138] It is particularly pointed out that due to the need of the use of the iteration mechanism, a plurality of order generators and classifiers should be configured, and there are two schemes: one is that each order can be the same generator and classifier; the second is that in the pre-training stage, the generators and classifiers of each order are independently trained according to different fine-grained levels, although this scheme will increase the consumption of computing power, but can further improve the accuracy of generation and classification of each order. When designing the model, the specific scheme should be selected according to the actual needs.
[0139] Through the above steps, the mel-spectrogram Mel-BC based on bone conduction signal and the mel-spectrogram Mel-Vba based on the collaborative coding of the generative adversarial network and the cross-modal attention mechanism are obtained. Since the mel-spectrograms of the two channels have their own advantages in representing speech information, it is still necessary to further fuse the mel-spectrograms of the two channels through step S33 to obtain the final mel-spectrogram Mel-U1t input to the vocal synthesis model, so that the description of the original speech is comprehensive and true.
[0140] Step S33: Fuse the mel-spectrogram Mel-Vba based on modal fusion with the mel-spectrogram Mel-BC based on the original bone conduction signal obtained in step S21 to obtain the final mel-spectrogram Mel-Ult.
[0141] The fusion of the mel-spectrogram is divided into two cases: silent scene and high noise scene, which will be described in detail below.
[0142] In the silent scene, the situation is relatively simple, the person does not produce a considerable sound excitation, and the signal richness perceived by the bone conduction sensor is low, so the mel-spectrogram Mel-BC based on bone conduction signal is ignored, and the Mel-Vba is directly used as the final mel-spectrogram Mel-Ult in the full frequency band, that is:
[0143] Mel-Ult = 1 * Mel-Vba + 0 * Mel-BC
[0144] In high noise scenes, that is, people do produce quite sound excitation, but the air conduction voice is invalid due to noise interference. In this case, the embodiment adopts a time-frequency partition image fusion method for the mel-spectrogram, and the core idea is to divide the mel-spectrogram in local areas according to the essential characteristics of the mel-spectrogram from the time domain and the frequency domain, so that each local area of the final mel-spectrogram Mel-Ult retains the optimal representation from the corresponding local area of Mel-BC and Mel-Vba. For the area division of the mel-spectrogram, see Figure 4 .
[0145] The horizontal axis of the mel-spectrogram is time, the vertical axis is frequency, and the depth of color represents the strength (amplitude) of the corresponding frequency component at a certain time. For easy calculation, it is converted into a gray image, and its amplitude can be described by a one-dimensional linear value P from 0 (black) to 255 (white). Since white has the highest energy and black has the lowest energy, in order to facilitate the direct correspondence between the numerical value and the energy size, the amplitude of a certain frequency at a certain time is recorded as:
[0146] A(t, h) = P
[0147] First, the partition and fusion are performed from the frequency dimension, and the purpose is to determine the high frequency band of the final mel-spectrogram. First, a cutoff frequency z is given, and the mel-spectrogram is divided into a high frequency partition H0 and a low frequency partition L according to the cutoff frequency z. This is because the bone conduction signal has a high frequency component that decays severely and disappears, and it is generally believed that the component above the cutoff frequency completely disappears. Therefore, the content of the high frequency partition above the cutoff frequency needs to be filled with the mel-spectrogram Mel-Vba based on the generative adversarial network and the cross-modal attention mechanism collaborative coding. And the cutoff frequency is also strongly related to the human body medium through which the bone conduction signal is transmitted, that is, it can be considered that after the placement position of the bone conduction sensor is determined, the cutoff frequency can also be determined accordingly. In this embodiment, the bone conduction sensor is placed on the cheek, and the cutoff frequency is about 3KHZ, that is, z = 3:
[0148] Mel(H0) = Mel(Clip | FRQ > 3KHZ)
[0149] Mel(L) = Mel(Clip | 0 < FRQ <= 3KHZ)
[0150] The part of Mel-BC in the high frequency partition H0 is removed and filled with the high frequency partition H0 of the mel-spectrogram Mel-Vba; thus, the high frequency partition of the final mel-spectrogram can be determined as:
[0151] Mel-Ult(H0) = 1 * Mel-Vba(H0) + 0 * Mel-BC(H0)
[0152] Secondly, partitioning and fusing from the time domain dimension, the purpose is to determine the middle and low frequency band of the final state mel-spectrogram. In the middle and low frequency band, the bone conduction signal has a more accurate time-frequency distribution, but the amplitude is slightly attenuated. At the same time, since the original sound excitation does not pass through the oral cavity, interlabial, nasal cavity and other regions, the consonant syllable relying on the friction and explosion of such regions will be lost. Therefore, what needs to be done is to reasonably enhance the amplitude of the corresponding region on the basis of the time-frequency distribution of the middle and low frequency band of the bone conduction signal, and especially to restore the amplitude of the corresponding region of the consonant syllable.
[0153] For amplitude attenuation, the amplitude of the corresponding region needs to be reasonably enhanced. First, the amplitude distribution of the middle and low frequency region of Mel-BC is calculated, and it is considered that:
[0154] In Mel-BC(L), if the amplitude at a certain time and a certain frequency is greater than a certain threshold x, but less than a certain threshold y (that is, within a certain interval), it is determined that this is an accurate time-frequency distribution point of audio information, but its amplitude has a certain attenuation, then the amplitude of the same time-frequency point of Mel-Vba(L) is reasonably enhanced; otherwise, no enhancement is done; in this embodiment, x = 30, y = 80, so the amplitude distribution of the middle and low frequency region of the final state mel-spectrogram can be determined as:
[0155] A(t, h) | Mel-Ult(L) = 1 * A(t, h) | Mel-Vba(L) (30 < A(t, h) | Mel-BC < 80)
[0156] A(t, h) | Mel-Ult(L) = 1 * A(t, h) | Mel-BC(L) (other)
[0157] On this basis, for the case of consonant syllable loss, the middle and low frequency region needs to be further partitioned into n equal time length small regions from the time domain dimension according to the time resolution, such as Figure 5 as shown in
[0158]
[0159]
[0160] Since each utterance is a command sentence, it can be considered that the Mel-BC(L) spectrogram concerned should have an uninterrupted amplitude distribution in the time domain. Therefore, when examining any small region, if the small regions before and after it have a certain amplitude distribution, the region should also have a certain amplitude distribution. If the region has no amplitude distribution, it is determined that the consonant syllable is missing, and the corresponding region of Mel-Vba(L) should be used to supplement. The specific process is as follows: first, the amplitude of each small region of Mel-BC(L) is integrated to calculate the amplitude sum A(L i ) in the small region Mel-BC(L i ), the time domain upper and lower limits are T i0 and T ie , and the frequency domain upper and lower limits are 0 to the cutoff frequency z, i.e. 0 to 3KHZ:
[0161]
[0162] When A(L i-1 ) and A(L i+1 ) are greater than a certain threshold p (only A(L i+1 ) is considered for the first region, and only A(L i-1 ) is considered for the tail region), and A(L i ) is less than a certain threshold q, the corresponding small region is filled with Mel-Vba(L):
[0163] Mel-Ult(L i )=1*Mel-Vba(L i ) (A(L i )<q&&A(L i-1 )>p&&A(L i+1 )>p)
[0164] Mel-Ult(L i )=1*Mel-BC(L i ) (other)
[0165] Therefore, the middle and low frequency partition of the final state Mel spectrogram in the high noise scene can be determined as:
[0166]
[0167] Therefore, the final state Mel spectrogram in the high noise scene can be determined as:
[0168] Mel-Ult=Mel-Ult(H0)+Mel-Ult(L).
[0169] Thus, the final state Mel spectrogram Mel-Ult based on Mel spectrogram fusion in the silent scene and the high noise scene is obtained.
[0170] The above illustrates the time-frequency partition image fusion method for the mel-spectrogram provided by the present application, and the overall flow is as shown in Figure 6 The mel-spectrogram is divided into regions from two dimensions of time domain and frequency domain, and image fusion is performed according to the time-frequency distribution characteristics of the modalities to obtain the original modality information, and more comprehensive information coding is obtained in the presentation form of the mel-spectrogram. In the present application, the single-modality mel-spectrogram based on the bone conduction signal and the mel-spectrogram fused from the two modalities of the bone conduction signal and the lip movement image signal are spliced and fused at the image level. In other application examples, the above method can also be extended to the corresponding processing of the mel-spectrogram derived from other information modalities.
[0171] Step S4 mainly includes:
[0172] Based on the obtained modality collaborative feature expression, i.e., the mel-spectrogram Mel-Ult, a pre-trained back-end classification neural network model can be used to correspondingly recognize the pre-training target, which is a specific short sentence, a text instruction, a logic instruction, etc. in the present embodiment.
[0173] As shown in Figure 6 Based on the final-state mel-spectrogram Mel-Ult after image fusion, a post-processing network is used to convert it into a linear frequency spectrum. The post-processing network adopts a mature 1-D Convolution Bank+Highway Network+Bidirectional GRU architecture. The obtained linear frequency spectrum is input into a mature vocoder based on the Griffin-Lim algorithm, i.e., it can be converted into a speech waveform to realize speech synthesis.
[0174] The speech synthesis of the present application introduces the bone conduction signal, which contains the pronunciation characteristics (tone, rhythm, timbre, etc.) of the speaker, so the synthesized speech has high restoration degree to the speaker; at the same time, by calling the pre-training model of the registered user during synthesis, the synthesized speech can further retain the timbre of the speaker himself.
[0175] Optionally, in another embodiment, after obtaining the single-modal data features determined based on the bone conduction signal and the lip movement image signal in the time domain and the spatial domain, the feature fusion and weight distribution within and between modalities based on the multi-head attention mechanism (Multi-head attention) are performed through the encoding and decoding of the Transformer model, and the multi-modal information based on the feature fusion is constructed; based on the language classification model, the multi-modal information is mapped to text information, and then the synthesized audio information is obtained through the TTS (Text to speech) text-to-speech conversion model such as Tacotron. Since the text result loses the original audio characteristics, if further preservation of the sound characteristics, tone, rhythm, timbre, etc. of the speaker in the synthesized speech is required, the feature expression of the bone conduction signal needs to be integrated into the existing TTS model.
[0176] Optionally, in another embodiment, after obtaining the single-modal data features determined based on the bone conduction signal and the lip movement image signal in the time domain and the spatial domain, the feature fusion and weight distribution within and between modalities based on the multi-head attention mechanism (Multi-head attention) are performed through the encoding and decoding of the Transformer model, and the multi-modal information based on the feature fusion is constructed; based on the language classification model, the multi-modal information is mapped to text information, and then the synthesized audio information is obtained through the TTS (Text to speech) text-to-speech conversion model such as Tacotron. Since the text result loses the original audio characteristics, if further preservation of the sound characteristics, tone, rhythm, timbre, etc. of the speaker in the synthesized speech is required, the feature expression of the bone conduction signal needs to be integrated into the existing TTS model.
[0177] The core architecture of bone conduction-lip reading fusion-synthesis includes a data acquisition unit, a feature extraction unit, an encoding unit, a speech synthesis unit, and an interaction unit:
[0178] Through the data acquisition unit, the bone conduction speech signal and the lip movement image signal synchronously acquired during user speech input are sent to the feature extraction unit;
[0179] Through the feature extraction unit, the received bone conduction speech signal and lip movement image signal data are preprocessed and feature-extracted respectively to determine the single-modal data features in the time domain and the spatial domain, which are sent to the encoding unit;
[0180] Through the encoding unit, based on the received single-modal data features in the time domain and the spatial domain, the generative adversarial network with the cross-modal attention mechanism and the mel-spectrogram fusion method are applied to establish a speech model and obtain the modal collaborative feature expression, which is sent to the speech synthesis unit;
[0181] Through the speech synthesis unit, according to the modal collaborative feature expression, using the post-processing network and vocoder, and combining the timbre, rhythm, pause habits and other features obtained from the user example audio during the previous user registration, a speech waveform with personalized timbre is synthesized and sent to the interaction unit;
[0182] Through the interaction unit, the synthesized speech result is evaluated in quality, and subsequent transmission is performed through the existing communication channel. The quality evaluation includes: objective index evaluation, calculating the ESTOI (Extended short-time objective intelligibility) and PESQ (Perceptual evaluation of speech quality) of the generated speech waveform to evaluate its intelligibility and perceptual quality, and scoring below a certain threshold is considered as unusable speech; subjective index evaluation, feedback to the speaker himself, and after perception, confirm whether the information is accurate and the audio is clear, and in the application scenario, a selection button can be set for human intervention to confirm whether to use the audio output.
[0183] The role of the visual information processing unit is to process the face image and scene image obtained based on the perception module, and finally realize high-quality video call effect and video interaction. The processing flow of visual information is as shown in Figure 7 The specific implementation method of the video interaction function is as follows:
[0184] The visual information processing unit includes a conventional computing unit and a three-dimensional immersive computing unit. The visual information processing unit is divided into a conventional mode and a three-dimensional immersive mode. Users can switch between the two modes through physical buttons according to scene requirements.
[0185] When switched to the conventional mode, the three-dimensional immersive mode computing unit enters a dormant state to reduce power consumption, and the system realizes conventional video calls based on two-dimensional images. Due to the process accuracy or assembly process of the camera, distortion of the original image may occur, that is, the coordinate position of each pixel point after imaging is inconsistent with the ideal projected coordinate position. For this reason, the conventional computing unit corrects the distortion of the face image to improve the user's perception in video calls and also ensures the authenticity and accuracy of the transmitted information. In this embodiment, the face image is captured by a close-up camera, and the main distortion form introduced by it is barrel distortion in radial distortion, that is, a distortion form distributed along the lens radius, the distortion rate at the optical axis center is 0, and the distortion gradually increases as the lens radius moves to the edge; and the magnification at the optical axis center is greater than that at the edge. The ideal pixel point coordinates (x, y) and the distorted pixel point coordinates (x d , y d ) have the following relationship:
[0186] x d= x + x[k1(x 2 + y 2 )+ k2(x 2 + y 2 ) 2 ]
[0187] y d = y + y[k1(x 2 + y 2 )+ k2(x 2 + y 2 ) 2 ]
[0188] where k1, k2 are radial distortion coefficients, after using Zhang Zhengyou calibration method to calculate the radial distortion coefficients, according to the relationship of the above formula, the ideal pixel point (x, y) can be solved from the distorted pixel point by inverse distortion, that is, the corrected image can be obtained; the image stream in the original video is corrected according to the above process, and the video output in the conventional mode can be obtained. The video output and the scene image obtained by the second visual perception unit are sent to the communication module and transmitted to the opposite terminal, and the video interaction function in the conventional mode can be realized.
[0189] In the three-dimensional immersion mode, the conventional calculation unit still works, and the image stream is first corrected for distortion according to the above method, and then processed by the three-dimensional immersion calculation unit, so as to realize real-time three-dimensional reconstruction of the portrait and three-dimensional immersive video call.
[0190] The implementation of the three-dimensional immersion mode adopts the classic three-dimensional deformable face model 3DMM (3D Morphable models) method in three-dimensional face modeling, and the basic idea is to regard the face model in three-dimensional space as a weighted combination of a set of orthogonal face basis vectors. In this embodiment, a three-dimensional face model based on the basic features of the user's face and combining real-time facial expressions is constructed through the two-dimensional face image obtained by the visual perception unit, and a set of orthogonal face basis vectors is established based on the user's trusted face data. The specific three-dimensional immersion calculation includes:
[0191] First, a structured light system is used to perform three-dimensional scanning on the user's face, and by adjusting the lighting conditions, the user's static and expressionless three-dimensional contour features in a certain light intensity interval and light angle interval are obtained, and (u, v) represents different intensity intervals and illumination angles.
[0192] V (u,v) = {X1, Y1, Z1, X2, Y2, Z2, …, X n , Y n , Z n} ∈ R 3n
[0193] The mapping relationship Zi =αXi+βY i i = 1, 2, ..., n
[0194] Simultaneously, acquire two-dimensional facial feature point data of users under different basic expression categories (including anger, disgust, fear, sadness, anticipation, happiness, surprise, trust, etc.). j j is the subscript for the emoji category;
[0195] V j ={x1,y1,x2,y2,…,x n y n}∈R 2n
[0196] Extended 3D facial feature data can be obtained as V j-e ;
[0197] z i =αx i +βy i i = 1, 2, ..., n
[0198] V j-e ={x1,y1,z1,x2,y2,z2,…,x n y n , z n}∈R 3n
[0199] The static, expressionless 3D contour features are randomly combined with extended 3D facial feature data to form a shape basis vector S. The combination principle is inspired by Dropout, as detailed below:
[0200] r (j-e) ~Bernoulli(p)
[0201] V j-e ′=r (j-e) *V j-e
[0202] S (u,v) =V (u,v) +W j-e V j-e
[0203] (The wavy lines above the symbols in the formula have no special meaning and are only used to distinguish them from the previous text);
[0204] r is a set of n 0s or 1s randomly generated by a Bernoulli distribution with probability p. This is then followed by V. j-eMultiply the vectors to randomly discard or retain a portion of the extended 3D facial feature vectors, then multiply by the weights of each extended 3D facial feature vector, and finally add the static, expressionless 3D contour features; W j-e These are the weights of each set of three-dimensional vectors;
[0205] Simultaneously, a 3D scan of the user's face is performed to obtain the texture features T of the user in a static, expressionless state. (u,v) ,
[0206] T (u.v) ={R1,G1,B1,R2,G2,B2,…,R n G n B n}∈R 3n
[0207] R, G, and B represent color components;
[0208] The above data matrix S (u,v) and T (u, v ) Principal component analysis (PCA) was used for dimensionality reduction to obtain two PCA models for shape features and texture features, respectively:
[0209]
[0210]
[0211] These are the means of two shape features and one texture feature, respectively.
[0212] V S =[v s1 v s2 v s3 , ..., v sm ]∈R 3n*m V T =[v t1 v t2 v t3 , ..., v tm ]∈R 3n*m V S V T They are shape features S (u,v ) and texture features T (u,v) m principal components obtained through principal component analysis, σ∈R m The standard deviation is represented by the standard deviation. Based on this, the 3D face model, including the shape model and the color model, can be obtained as follows. By superimposing the two models, the final modeling result I (3D reconstruction model) can be obtained.
[0213] Shape model:
[0214] Texture model:
[0215] Three-dimensional reconstruction model: I = S + T
[0216] wherein λ i and ρ i represent shape parameters and texture parameters, respectively;
[0217] The shape and texture principal component analysis model M S , M T of the obtained base vector group is stored after being labeled according to the scanned user, to obtain a face base model database;
[0218] In the three-dimensional immersion mode, the key feature parameters of the two-dimensional face image acquired by the first visual perception unit are identified in real time by calling the corresponding face base model data, and the selection of the group of feature parameters is consistent with the two-dimensional face feature point data V j of the user established in the foregoing, and based on the prior model, the corresponding shape parameters λ i are obtained.
[0219] On the other hand, the key feature points affected by the lighting condition are additionally selected, and the bilinear interpolation method is used to solve the texture parameters ρ i constrained by the lighting condition.
[0220] Based on the reliable face base vector group, i.e., the real-time acquired face base vector group (including the shape base vector S and the texture (color) base vector T, which can be regarded as the initial S and T), and the shape parameters and the texture parameters, the real-time mapping of the two-dimensional face image to the three-dimensional face image is realized by using the foregoing three-dimensional reconstruction model. The reconstructed three-dimensional face image stream and the current environment image sequence acquired by the second visual perception unit are sent to the communication module, transmitted to other terminals, and presented, so that the video interaction function in the three-dimensional immersion mode is realized.
[0221] The processing flow of the physiological information is shown in Figure 8 ; and the working method of the physiological information processing unit is as follows:
[0222] The physiological information processing unit measures the heart rate of the user by using the remote photoplethysmographic (rPPG) method based on the face image of the user acquired by the first visual perception unit; and the heart rate measurement result is transmitted to the display module.
[0223] The remote photoplethysmographic (rPPG) method refers to a technology of capturing the periodic change of the skin color caused by the heart cycle by using a camera or other sensors.
[0224] The blood flow caused by the beating of the heart forms a periodic change in the microvessels of the skin tissue of the human body, so that the absorption and reflection of light also has a periodic signal. The periodic signal change can be analyzed by capturing the face image through the camera to monitor the change of the heart rate. The specific method of the embodiment is to frame the specific region of interest ROI (Region of Interest) of the obtained face image; calculate the spatial mean of the RGB three channels in the framed region; use low-pass filtering, blind source separation, etc. on the three-channel spatial mean to obtain a component containing heart rate information; apply fast Fourier transform to the component to estimate the corresponding frequency F, and the heart rate can be calculated as 60*F. The heart rate calculation result is transmitted to the display module, and through the augmented reality glasses, the user can see the heart rate condition in real time; at the same time, the heart rate calculation result is sent to the communication module and then transmitted to other terminals, so that the other party or the host terminal can remotely monitor the heart rate condition of the user.
[0225] The physiological information processing unit processes the face image obtained by the first visual perception unit and the eye movement information obtained by the second visual perception unit;
[0226] For the face image, first, face alignment and normalization preprocessing are performed. On the one hand, attention is paid to the conventional 68 facial key feature points such as eyes, nose tip, corners of the mouth, eyebrows, etc. to realize the representation of the global basic information of the face; on the other hand, since micro-expression has an important role in the more true representation of emotion, attention is also paid to the micro-expression movements with a duration of only 1 / 25 to 1 / 5 seconds, such as eyebrow lowering, cheek rising, chin lowering, eyelid drooping, etc. which are not easy to be perceived, and attention is paid to the spatiotemporal domain features and optical flow features, etc. to obtain the dynamic representation of the information. Through the convolutional neural network CNN and the long short-term memory network LSTM, the extraction of the image features (micro-expression features) and the timing information is realized, and then the learned deep features are outputted by the shallow classifier (such as support vector machine SVM) to obtain the multi-classification result R0.
[0227] For the eye movement information, eye movement features are obtained, which include pupil diameter, gaze deviation, gaze duration, saccade duration, saccade amplitude, blink duration and blink frequency; wherein the principal component analysis is performed on the pupil diameter information, and the feature information is smoothed and normalized. Through the automatic encoder based on the restricted Boltzmann machine, the high-order feature expression is extracted from the eye movement features according to the weight, and the multi-classification result R1 is outputted by the shallow classifier (such as support vector machine SVM).
[0228] The multi-classification result is the different emotion categories recognized. The embodiment preliminarily pays attention to eight basic emotions such as anger, disgust, fear, sadness, anticipation, joy, surprise and trust, and can be extended to more fine-grained emotion categories.
[0229] The preliminary emotion recognition results R0, R1 are obtained by the above method; meanwhile, the features of the two channels are fused in the feature layer, that is, the micro-expression features and eye movement features of the foregoing face image are normalized and spliced, and then the results R2 are output by the rear-end classifier; the preliminary emotion recognition results R0, R1 and the fused emotion recognition result R2 are fused in the decision layer, the decision layer fusion is based on the Bayesian decision fusion method, based on the confidence of different classification results of R0, R1 and R2, the final preferred conclusion is calculated to obtain the final reliable emotion recognition result R.
[0230] The final reliable emotion recognition result R is sent to the display module and presented in a specific area of the user's field of view, so that the user can see the emotion recognition result of himself / herself, confirm whether it is accurate or not, and can be used as a reference. In this embodiment, considering the need to objectively let the communication partner or the host know the user's emotional state, the reliable emotion recognition result is sent to the communication module regardless of whether the user confirms the accuracy of the emotion recognition result or not, and then transmitted to the terminal of the other party, that is, the display and interaction function of physiological information. By designing the earphone interaction system, the integrated audio, video and physiological information multimedia communication interaction based on the head-mounted portable earphone is realized.
[0231] In summary, by using the above technical solutions, in daily environment and complex environment (high noise, high mobility), the present application can complete the recognition, synthesis and transmission of voice information through the perception and processing of multi-modal information such as audio, lip diagram, bone conduction, etc.; in addition, through the communication module and the small loudspeaker, the voice of the other party is received to realize high-quality real-time voice communication, so as to realize the information interaction between people and people, and between people and the host terminal by taking voice as the carrier. Through the processing of video images, real-time distortion correction of face images and real-time three-dimensional modeling of faces are completed, and the optimization, reconstruction and transmission of visual information are completed; in addition, through the communication module and the augmented reality glasses, the video of the other party is received to realize high-quality, three-dimensional immersive real-time video communication, so as to realize the information interaction between people and people, and between people and the host terminal by taking video as the carrier. Through the processing of video images, the features containing physiological information contained therein can also be extracted, the real-time heart rate and emotional categories of the user and other physical and mental indications can be calculated and obtained, the presentation of related physical and mental features based on the display device and the transmission based on the communication device are completed; in addition, through the communication module and the augmented reality glasses, the physiological indications of the other party are received to realize accurate and convenient physical and mental state monitoring, so as to realize the information interaction between people and people, and between people and the host terminal by taking physiological indications as the carrier.
[0232] The present application also provides an interactive earphone, which comprises the above-mentioned interactive earphone system; that is, it comprises a microphone, a bone conduction sensor, a face camera, a loudspeaker and augmented reality glasses.
[0233] The microphone is placed at a certain distance from the user's mouth by fixing the adjustable microphone rod, and perceives the air conduction audio when the user speaks. The bone conduction audio sensor is attached to the surface skin of the user's cheek, and perceives the bone conduction signal generated when the user speaks in a high noise environment.
[0234] The face camera is fixed by the adjustable microphone rod, and is placed at a certain distance from the user's face. The face camera serves as a first visual information perception unit, and captures the face image of the user in real time. The face camera can simultaneously include a local image of the mouth area for lip reading, a global image for video call and physical and mental monitoring, and a specific area local image, and the corresponding frame is taken according to the processing requirement.
[0235] The outer side of the augmented reality glasses is provided with a scene camera, which serves as a second visual information perception unit and captures the scene image in front of the user in real time. The inner side of the augmented reality glasses is provided with a macro camera, which serves as a third visual information perception unit and captures the eye movement information of the user in real time.
[0236] The above is only the preferred embodiment of the present application, and does not limit the present application in any way. Any person skilled in the art can make any form of equivalent replacement, modification or change to the technical solutions and technical contents disclosed in the present application without departing from the scope of the technical solutions of the present application, and such change still falls within the protection scope of the present application.
Claims
1. An interactive headphone system, characterized in that, include: A perception module is used to acquire information from various modalities; the perception module includes an audio information perception unit and a visual information perception unit; the audio information perception unit of the perception module includes a microphone and a bone conduction signal sensor; the visual information perception unit includes a first visual perception unit, a second visual perception unit, and a third visual perception unit; the first visual perception unit acquires a user's facial image, the second visual perception unit acquires a scene image, and the third visual perception unit acquires a user's eye movement image; The microphone acquires air conduction audio signals, the first visual information perception unit acquires lip image signals, and the bone conduction signal sensor acquires bone conduction signals. The acquired signals are processed by the audio information processing unit to output specific phrases and instructions, and synthesize audio. The user's facial image is acquired by the first visual perception unit, and the scene image is acquired by the second visual perception unit. The acquired facial image and scene image are processed by the visual information processing unit to realize conventional video calls or three-dimensional immersive video calls. The user's facial image is obtained by the first visual perception unit, and the eye movement image is obtained by the third visual perception unit. The obtained facial image and eye movement image are processed by the physiological information processing unit to obtain the user's heart rate and the user's emotion type. The computation module is used to process the information of the various modalities and obtain the processing results; the computation module includes an audio information processing unit, a visual information processing unit, and a physiological information processing unit; The communication module is used to transmit the processing results of the computing module to other terminals; receive information from other terminals and send it to the display module. The display module is used to present auditory and visual information.
2. The interactive earphone system according to claim 1, characterized in that, The audio information processing unit uses air-conducted audio as the target mode for sound pickup, converts and transmits the air-conducted audio, and realizes audio interaction. Alternatively, the audio information processing unit can use lip image signals and bone conduction signals as the target modalities for sound pickup, and achieve modal information fusion audio interaction by using the bone conduction-lip reading fusion-synthesis method; The bone-guided lip-reading fusion-synthesis method includes: Bone conduction speech signals and lip movement image signals are acquired synchronously during user voice input; Based on the bone conduction speech signal and lip movement image signal, determine the single-modal data features in the time domain and spatial domain; Based on the determined single-modal data features in the time and spatial domains, a generative adversarial network incorporating a cross-modal attention mechanism and a Mel spectrogram fusion method are applied to obtain modal collaborative feature representations. Based on the obtained modal collaborative feature representation, a trained backend classification neural network model is applied to output specific phrases and instructions; a human voice synthesis model is applied to obtain audio waveforms.
3. The interactive earphone system according to claim 1, characterized in that, The visual information processing unit includes a conventional computing unit and a three-dimensional immersive computing unit. The computing of the visual information processing unit is divided into a conventional mode and a three-dimensional immersive mode. In normal mode, the conventional computing unit corrects distortions in the facial image, using ideal pixel coordinates (x, y) and distorted pixel coordinates (x, y). d y d The following relationship exists: x d =x+x[k1(x 2 +y 2 )+k2(x 2 +y 2 ) 2 ] and d =y+y[k1(x 2 +and 2 )+k2(x 2 +and 2 ) 2 ] Where k1 and k2 are radial distortion coefficients, after calculating the radial distortion coefficients using Zhang Zhengyou's calibration method, based on the relationship in the above formula, the distortion of the pixel (x) is calculated through inverse distortion. d y d Solve for the ideal pixel point (x, y) to obtain the corrected image and realize video calling based on the two-dimensional image; In 3D immersive mode, the conventional computing unit still works, first correcting the distortion of the image, and then processing the corrected image through the 3D immersive computing unit to reconstruct the portrait in real time, thus realizing 3D immersive video calls.
4. The interactive earphone system according to claim 3, characterized in that, The calculation of the three-dimensional immersive mode includes: A structured light system is used to perform a three-dimensional scan of the user's face. By adjusting the lighting conditions, the three-dimensional contour features of the user's static, expressionless face are obtained under a certain range of light intensity and light angle. (u, v) represent different intensity ranges and illumination angles. Regression mapping relationship Simultaneously acquire two-dimensional facial feature point data of users under different expression categories. j is the subscript for the emoji category; Extended 3D facial feature data can be obtained as V j-e ; The static, expressionless 3D contour features are randomly combined with the extended 3D facial feature data to form a shape basis vector S, as follows: Simultaneously, a 3D scan of the user's face is performed to obtain the texture features T of the user in a static, expressionless state. (u,v) , The above data matrix S (u,v) and T (u,v) Principal component analysis (PCA) was used for dimensionality reduction to obtain two PCA models for shape features and texture features, respectively: These are the means of two shape features and one texture feature, respectively. They are shape features S (u,v) and texture features T (u,v) m principal components obtained through principal component analysis, σ∈R m The standard deviation is represented by the standard deviation. Based on this, the 3D face model includes a shape model and a color model as follows. The two models are superimposed to obtain the final modeling result I. Shape model: Texture model: 3D reconstruction model: I = S + T in and These represent shape parameters and texture parameters, respectively. The obtained principal component analysis model M S M T After labeling and storing the scanned users, a basic facial model database is obtained; In 3D immersive mode, the corresponding basic facial model data is invoked to identify key feature parameters of the 2D facial image acquired by the visual perception unit in real time. The selection of this set of feature parameters is related to the aforementioned establishment of the user's 2D facial feature point data. Consistent with the above prior model, the corresponding shape parameters can be obtained. The aforementioned prior models include two principal component analysis models, a shape model, a texture model, and a 3D reconstruction model; Next, key feature points affected by lighting conditions are selected, and bilinear interpolation is used to solve for the texture parameters constrained by lighting conditions. ; Based on the real-time acquired facial base vector set, as well as shape and texture parameters, the aforementioned 3D reconstruction model is used to map the 2D facial image into a 3D facial image in real time.
5. An interactive headphone system according to claim 1, characterized in that, The physiological information processing unit measures the user's heart rate using remote photoplethysmography based on the user's facial image acquired by the first visual perception unit; and transmits the heart rate measurement result to the display module.
6. The interactive earphone system according to claim 1, characterized in that, The physiological information processing unit processes the facial image acquired by the first visual perception unit and the eye movement information acquired by the second visual perception unit. For facial images, face alignment and normalization preprocessing are performed first. Then, facial image features and temporal information are extracted using a convolutional neural network (CNN) and a long short-term memory network (LSTM). Finally, the obtained deep features are processed by a shallow classifier to output multi-classification results. The facial image features include facial feature points, as well as spatiotemporal features and optical flow features of facial micro-expressions. For eye movement information, eye movement features are acquired, including pupil diameter, fixation deviation, fixation duration, saccade duration, saccade amplitude, blink duration, and blink frequency. Principal component analysis is performed on the pupil diameter information, and the features are smoothed and normalized. An autoencoder based on a restricted Boltzmann machine is used to encode various eye movement features according to weights, extract high-order feature representations, and output multi-classification results R1 through a shallow classifier. Simultaneously, facial image features and eye movement features are normalized and stitched together, and then the emotion recognition result is output by the back-end classifier. ; preliminary emotion recognition results , And the fusion of emotion recognition results Decision-level fusion is performed to obtain the final emotion recognition result R.
7. An interactive headset, characterized in that, The interactive headset system includes any one of claims 1-6; the interactive headset system includes a microphone, a bone conduction sensor, a facial camera, a speaker, and augmented reality glasses; The augmented reality glasses have a scene camera on the outside and a macro camera on the inside.
8. An interactive headset according to claim 7, characterized in that, The facial camera serves as the first visual information sensing unit, capturing the user's facial image in real time; the scene camera serves as the second visual information sensing unit, capturing scene images in real time; and the macro camera serves as the third visual information sensing unit, capturing the user's eye movement information in real time.
Citation Information
Patent Citations
Complex scene voice recognition method and device based on multiple modes
CN112151030A
Multi-modal real-time emotion recognition method and system based on DS evidence theory
CN114463827A