Speech processing method, apparatus, and XR device
By acquiring scene images and head posture data in XR devices, and combining real-time audio and video data processing from built-in cameras and microphones, visual and audio information is fused, solving the problem of accurately separating target speech in noisy environments and improving the interactive experience and convenience.
Patent Information
- Application Number
- CN202511049200.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Existing technologies struggle to accurately identify target speech signals in XR devices without increasing hardware costs or relying on prior information about the target speaker, especially in noisy environments where multiple people are conversing and it is difficult to effectively separate the speech of a specific speaker.
By acquiring current scene images and user head posture data, the target object being gazed at is determined. Real-time audio and video data are acquired using the XR device's built-in camera and microphone, and then processed using a speech separation model to fuse visual and audio information to accurately separate the target audio signal.
It achieves accurate separation of the target speaker's voice in noisy environments, enhancing the immersiveness and convenience of the interactive experience. It has low hardware requirements, requires no additional expensive hardware, and is suitable for existing XR devices.
Smart Images

Figure CN120564748B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech processing, and in particular to a speech processing method and device and an XR device. BACKGROUND
[0002] In a multi-person conversation or a noisy environment (i.e., a "cocktail party" scenario), the "cocktail party effect" exhibited by the human auditory system enables it to highly focus its attention on a specific speaker object, effectively suppressing the interference of background noise and other speaker objects. However, for machines, especially on lightweight devices such as XR (Extended Reality) devices, it is always a challenging task to automatically and accurately separate the speech signal of a specific target speaker object from the mixed audio collected by a single microphone.
[0003] To address this challenge, traditional technical solutions often rely on additional hardware devices or prior information of the target speaker object. One common solution is to use a microphone array to enhance the sound of the target speaker object in the spatial direction through beamforming. However, a microphone array significantly increases the cost, size, and complexity of the device. Another solution uses the voiceprint features of the target speaker object for speech separation. For example, by pre-acquiring an audio sample of the target speaker object to extract its voiceprint features, the speech signal of the target speaker object can be identified and extracted from the mixed speech. Although this selective listening method based on voiceprint features eliminates the dependence on the location of the speaker object, its core defect is that it must have an effective voiceprint sample of the target speaker object in advance. This significantly limits its applicability in common application scenarios such as temporary conversations or first meetings.
[0004] In summary, although existing technologies have shown the potential to selectively extract the speech of a target speaker object under certain conditions, existing solutions are difficult to apply to XR devices, which cannot meet the flexible and convenient practical application requirements while maintaining the portability and low cost of XR devices. Therefore, how to accurately identify the speech of a target speaker object without increasing the additional hardware cost of an XR device and without relying on prior information of the target speaker object is a key technical problem that needs to be solved. SUMMARY
[0005] The present application provides a speech processing method, device and XR device to accurately identify the speech signal of a target speaker object without increasing the additional hardware cost of an XR device and without relying on prior information of the target speaker object.
[0006] The present application provides a speech processing method applied to an XR device, comprising:
[0007] obtain a current scene image and current head posture data of a user;
[0008] determine a target gaze object according to the current scene image and the current head posture data;
[0009] obtain real-time audio data and real-time video data of the target gaze object;
[0010] process the real-time audio data and the real-time video data through a speech separation model to determine a target audio signal of the target gaze object.
[0011] According to the speech processing method provided by the application, after determining the target gaze object according to the current scene image and the current head posture data, the method further comprises:
[0012] obtain a first head posture sequence data of the user, and determine a current reference trajectory according to the first head posture sequence data;
[0013] obtain a second head posture sequence data of the user in real time, and detect whether the gaze object of the user changes according to the second head posture sequence data and the current reference trajectory;
[0014] if it is detected that the gaze object of the user does not change, continue to obtain the real-time audio data and the real-time video data of the target gaze object;
[0015] if it is detected that the gaze object of the user changes, re-obtain the current scene image and the current head posture data of the user.
[0016] According to the speech processing method provided by the application, the second head posture sequence data of the user is obtained in real time, and whether the gaze object of the user changes is detected according to the second head posture sequence data and the current reference trajectory, which comprises:
[0017] obtain a second head posture sequence data of the user in real time, and determine a real-time motion trajectory according to the second head posture sequence data;
[0018] encode the real-time motion trajectory into a real-time feature vector, and encode the current reference trajectory into a current reference feature vector;
[0019] calculate a deviation degree according to the real-time feature vector and the current reference feature vector;
[0020] detect whether the deviation degree is greater than a preset deviation threshold to detect whether the gaze object of the user changes.
[0021] The voice processing method provided in the application comprises the following steps:
[0022] Detecting whether the deviation degree is greater than a preset deviation degree threshold value;
[0023] If it is detected that the deviation degree is less than or equal to the preset deviation degree threshold value, it is determined that the gaze object of the user does not change;
[0024] If it is detected that the deviation degree is greater than the preset deviation degree threshold value, a timer is started, and it is monitored whether the deviation degrees in a preset time are all greater than the preset deviation degree threshold value;
[0025] If it is monitored that the deviation degrees in the preset time are all greater than the preset deviation degree threshold value, it is determined that the gaze object of the user changes.
[0026] The voice processing method provided in the application comprises the following steps:
[0027] The real-time audio data is processed by a time sliding window, and the real-time video data is processed by a frame;
[0028] The real-time audio data processed by the time sliding window and the real-time video data processed by the frame are input into the voice separation model, so as to obtain the block audio signal corresponding to the target gaze object output by the voice separation model;
[0029] The block audio signal is spliced by a smoothing window, so as to obtain the target audio signal of the target gaze object.
[0030] The voice processing method provided in the application comprises the following steps:
[0031] The real-time audio data processed by the time sliding window is input into the audio encoder, so as to obtain the audio feature output by the audio encoder;
[0032] The real-time video data processed by the frame is input into the visual encoder, so as to obtain the face visual feature output by the visual encoder;
[0033] input the audio feature and the face visual feature into the audio-video fusion module to obtain an audio-video fusion feature output by the audio-video fusion module;
[0034] input the audio-video fusion feature into the mask estimation network to obtain a time-frequency mask output by the mask estimation network;
[0035] input the time-frequency mask into the audio decoder to obtain a block audio signal corresponding to the target gaze object output by the audio decoder.
[0036] According to the speech processing method provided by the application, the target gaze object is determined according to the current scene image and the current head posture data, and the method comprises the steps that:
[0037] performing face detection on the current scene image to obtain a candidate object and a pixel coordinate thereof;
[0038] establishing a coordinate system with the user's head as the origin, the front of the user's head as the X axis, and the vertical upward direction of the origin as the Z axis, and determining the Y axis according to the X axis and the Z axis;
[0039] determining the relative coordinate of the candidate object in the coordinate system according to the pixel coordinate and the current scene image;
[0040] calculating the face direction vector of the candidate object according to the relative coordinate and the reference coordinate of the user's head;
[0041] calculating the head orientation vector of the user according to the current head posture data;
[0042] calculating the included angle between the head orientation vector and the face direction vector, and determining the target gaze object from the candidate object according to the calculated included angle.
[0043] According to the speech processing method provided by the application, the target audio signal of the target gaze object is determined after the real-time audio data and the real-time video data are processed by the speech separation model, and the method further comprises the steps that:
[0044] performing intensity change processing on the target audio signal of the target gaze object in the real-time audio data; or
[0045] performing type conversion processing on the target audio signal of the target gaze object.
[0046] The application further provides a speech processing device comprising the following modules:
[0047] a first acquisition module configured to acquire a current scene image and current head posture data of a user;
[0048] an object determining module configured to determine a target gaze object according to the current scene image and the current head pose data;
[0049] a second obtaining module configured to obtain real-time audio data and real-time video data of the target gaze object;
[0050] an audio-video processing module configured to process the real-time audio data and the real-time video data by using a speech separation model to determine a target audio signal of the target gaze object.
[0051] The present application also provides an XR device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the speech processing method according to any one of the above when executing the computer program.
[0052] The speech processing method, device and XR device provided by the present application can obtain a current scene image and current head pose data of a user by using an XR device, so as to select and determine a target gaze object from the scene image seen by the user according to the orientation of the head of the user. The target object is determined in this way, which only needs to rely on the camera and IMU sensor provided by the XR device, and is more in line with the natural communication habits of human beings than manually selecting a target speaker object or relying on prior settings, thereby improving the immersion and convenience of the interactive experience. Then, the real-time audio data and real-time video data of the target gaze object are obtained by using the camera and microphone provided by the XR device, and the real-time audio data and real-time video data are processed by using a speech separation model to determine the target audio signal of the target gaze object. By fusing visual and audio information, the speech of the target speaker object can be accurately separated in a complex scene such as a noisy environment and multiple people speaking at the same time. In summary, the present application has low hardware requirements, and only a single camera and microphone provided by the XR device are needed to realize directional extraction of the speech of the target speaker object, without the need for additional expensive hardware such as a microphone array, a laser radar or an eye tracking device, and without the need for relying on prior information of the target speaker object, thereby facilitating integration and deployment on existing XR devices. BRIEF DESCRIPTION OF DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0054] Figure 1 is a system architecture diagram of the speech processing system provided by the present application;
[0055] Figure 2 is one of the flowcharts of the speech processing method provided by the present application;
[0056] Figure 3 is one of the flowcharts of the speech processing method provided by the present application;
[0057] Figure 4 is one of the flowcharts of the speech processing method provided by the present application;
[0058] Figure 5 is a structural schematic diagram of the speech processing device provided by the present application;
[0059] Figure 6 is a structural schematic diagram of the XR device provided by the present application. DETAILED DESCRIPTION
[0060] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0061] The present application proposes a speech processing method, device and XR equipment, which will be described below with reference to the drawings. Figures 1-6
[0062] As shown in Figure 1 , the speech processing system includes an extended reality (Extended Reality, XR for short) device 01 and a server 02. The XR device 01 includes but is not limited to a VR (Virtual Reality, VR for short) device, an AR (Augmented Reality, AR for short) device and an MR (Mixed Reality, MR for short) device, which is used to execute the speech processing method of the present application. The server 02 can be a server, which is used to train a speech separation model.
[0063] The server 02 deploys the trained speech separation model to the XR device 01 after training the speech separation model, and the XR device 01 can process real-time audio and video data (including real-time audio data and real-time video data) based on the speech separation model to determine the target audio signal of the target gaze object. Specifically, the user can wear the XR device 01, and by turning the head to gaze at the target speaker, the XR device 01 detects the head turning of the user, obtains the current scene image and the current head pose data of the user, and then determines the target gaze object according to the current scene image and the current head pose data. Further, the real-time audio data and the real-time video data of the target gaze object are obtained, and the real-time audio data and the real-time video data are processed by the speech separation model to determine the target audio signal of the target gaze object.
[0064] In some other embodiments of the present application, the XR device 01 can not deploy the speech separation model, and the server 02 processes the real-time audio and video data to determine the target audio signal of the target gaze object. Specifically, after the XR device 01 determines the target gaze object, the real-time audio and video data are obtained and sent to the server 02, the server 02 processes the real-time audio and video data by the speech separation model to determine the target audio signal, and then sends the target audio signal to the XR device 01, so that the XR device 01 can perform subsequent processing based on the target audio information, such as intensity change processing, type conversion processing, etc.
[0065] Figure 2 is one of the flowcharts of the speech processing method provided by the present application, as shown in Figure 2 The speech processing method includes steps S110, S120, S130 and S140.
[0066] Step S110, obtaining the current scene image and the current head pose data of the user.
[0067] In this embodiment, the speech processing method is applied to the XR device.
[0068] The XR device refers to a wearable or portable device that realizes human-computer interaction by fusing virtual and real environments through hardware and software technologies. The XR device includes, but is not limited to, a VR (Virtual Reality) device, an AR (Augmented Reality) device, and an MR (Mixed Reality) device. Among them, the VR device is a three-dimensional virtual space simulated by computer technology, which allows users to immerse in it and interact with it to obtain an immersive experience; the AR device is to fuse virtual information with the real world in real time, and superimpose it on the real scene to improve the sensory experience; the MR device is to mix the real world and the virtual world together to produce a new visual environment, which contains physical entities and virtual information, and users can interact with the physical entities and virtual information in real time.
[0069] The speech processing method of the present application is suitable for various scenes that need to extract specific speech from a noisy and multi-source environment, including but not limited to business meetings, social gatherings, medical consultations, education and training, security monitoring, etc.
[0070] In the present embodiment, the current scene image is an image obtained based on the current gaze direction of the user, which can be obtained through a camera of the XR device (such as a monocular RGB (Red-Green-Blue) camera), an RGB-D (RGB-Depth) camera or a binocular camera.
[0071] The head pose refers to the direction and angle of the user's head. The current head pose data of the user can be represented by a quaternion, which can be obtained by an IMU (Inertial Measurement Unit). The IMU is an inertial sensor module built-in the XR device, which generally includes an accelerometer and a gyroscope, and is used to measure the attitude angle (such as pitch, yaw, roll) and motion trajectory of the user's head in real time.
[0072] In step S120, a target gaze object is determined according to the current scene image and the current head pose data.
[0073] Here, the target gaze object is the speaking object that the user wants to focus on.
[0074] As an implementation manner, the current scene image is subjected to face detection to obtain a candidate object; if the candidate object is one, the candidate object is directly taken as the target gaze object. If the candidate object includes multiple objects, pixel coordinates of the candidate objects are obtained; a coordinate system is established with the user's head as the origin, the front direction of the user's head as the X axis, and the vertical upward direction of the origin as the Z axis, and the Y axis is determined according to the X axis and the Z axis; the relative coordinates of the candidate objects in the coordinate system are determined according to the pixel coordinates and the current scene image; the face direction vector of the candidate objects is calculated according to the relative coordinates and the reference coordinates of the user's head; the head orientation vector of the user is calculated according to the current head posture data; the included angle between the head orientation vector and the face direction vector is calculated, and the candidate object corresponding to the smallest included angle is selected as the target gaze object.
[0075] As another implementation manner, the current scene image is subjected to face detection to obtain a candidate object and pixel coordinates of the candidate object; a coordinate system is established with the user's head as the origin, the front direction of the user's head as the X axis, and the vertical upward direction of the origin as the Z axis, and the Y axis is determined according to the X axis and the Z axis; the relative coordinates of the candidate objects in the coordinate system are determined according to the pixel coordinates and the current scene image; the face direction vector of the candidate objects is calculated according to the relative coordinates and the reference coordinates of the user's head; the head orientation vector of the user is calculated according to the current head posture data; the included angle between the head orientation vector and the face direction vector is calculated, and the smallest included angle is selected from the calculated included angles. Further, it is judged whether the smallest included angle is less than or equal to a preset threshold value to determine whether the line of sight of the user's head has been aligned with the target object, wherein the preset threshold value can be a fixed value, or can be determined according to the standard deviation of the micro-motion of the user's head obtained through statistics. If the smallest included angle is less than or equal to the preset threshold value, the candidate object corresponding to the smallest included angle is determined as the target gaze object; if the smallest included angle is greater than the preset threshold value, it is determined that there is no target gaze object in the current scene image.
[0076] In step S130, real-time audio data and real-time video data of the target gaze object are obtained.
[0077] The real-time audio data can be obtained through the microphone of the XR, and the real-time video data can be obtained through the camera of the XR.
[0078] Further, a speech activity detection module can be integrated in the real-time audio stream to determine whether there is any speech of a speaking object in the obtained real-time audio data. When speech activity is detected, the real-time audio data and the real-time video data are further processed through a speech separation model.
[0079] At step S140, the real-time audio data and the real-time video data are processed by the speech separation model to determine the target audio signal of the target gaze object.
[0080] In view of the delay problem, when determining the target audio signal of the target gaze object by using the speech separation model, a streaming processing mode is adopted. Specifically, the real-time audio data is processed by a time sliding window for blocking, and the real-time video data is processed by a frame for blocking; then, the blocked real-time audio data and the blocked real-time video data are input into the speech separation model to obtain the blocked audio signal corresponding to the target gaze object output by the speech separation model; and then, the blocked audio signal is spliced by a smoothing window to obtain the target audio signal of the target gaze object.
[0081] Further, as an implementation, the speech separation model includes an audio encoder, a visual encoder, an audio-video fusion module, a mask estimation network, and an audio decoder. The processing process of the speech separation model on the blocked real-time audio data and the blocked real-time video data is as follows: the blocked real-time audio data is input into the audio encoder to obtain the audio feature output by the audio encoder, which is a time-frequency feature; at the same time, the blocked real-time video data is input into the visual encoder to obtain the facial visual feature output by the visual encoder; then, the audio feature and the facial visual feature are input into the audio-video fusion module to obtain the audio-video fusion feature output by the audio-video fusion module; the audio-video fusion feature is input into the mask estimation network to obtain the time-frequency mask output by the mask estimation network; and the time-frequency mask is input into the audio decoder to obtain the blocked audio signal corresponding to the target gaze object output by the audio decoder. The specific execution process can be referred to the following embodiments, which will not be described here.
[0082] Further, as another implementation, the speech separation model includes an audio encoder, a visual encoder, an audio-video fusion module, a time-domain convolution network, and an audio decoder. The processing process of the speech separation model on the blocked real-time audio data and the blocked real-time video data is as follows: the blocked real-time audio data is input into the audio encoder to obtain the audio feature output by the audio encoder, which is a time-domain feature; at the same time, the blocked real-time video data is input into the visual encoder to obtain the facial visual feature output by the visual encoder; then, the audio feature and the facial visual feature are input into the audio-video fusion module to obtain the audio-video fusion feature output by the audio-video fusion module; the audio-video fusion feature is input into the time-domain convolution network to obtain the time-domain mask output by the time-domain convolution network; and the time-domain mask is input into the audio decoder to obtain the blocked audio signal corresponding to the target gaze object output by the audio decoder.
[0083] Here, the speech separation model is trained based on sample audio-video data, sample face, and sample face audio signal annotation results.
[0084] The speech processing method provided by the application obtains the current scene image and the current head posture data of the user by the XR device, so as to select and determine the target gaze object from the scene image seen by the user according to the head orientation of the user. The target object is determined in the manner, which only needs to rely on the camera and the IMU sensor of the XR device, and compared with manually selecting the target speaking object or relying on prior settings, is more in line with the natural communication habits of human beings, and improves the immersion and convenience of the interactive experience. Then, the real-time audio data and the real-time video data of the target gaze object are obtained by the camera and the microphone of the XR device, and then the real-time audio data and the real-time video data are processed by the speech separation model to determine the target audio signal of the target gaze object. By fusing the visual and audio information, the speech of the target speaking object can be accurately separated in a complex scene such as a noisy environment and multiple people speaking at the same time. In summary, the hardware requirement of the application is low, and the directional extraction of the speech of the target speaking object can be realized by only a single camera and a microphone of the XR device, without the need for additional microphone arrays, laser radars or expensive hardware such as eye tracking, and without the need for relying on prior information of the target speaking object, which is convenient for integration and deployment on existing XR devices.
[0085] Figure 3 is a flowchart of the speech processing method provided by the application, as shown in Figure 3 After the step S120, the speech processing method further comprises steps S150 and S160.
[0086] In step S150, the first head posture sequence data of the user is obtained, and the current reference trajectory is determined according to the first head posture sequence data.
[0087] In this embodiment, it is considered that in a multi-person scene, the user may need to switch different speaking objects to obtain the speech of different objects. Therefore, in this embodiment, the current reference trajectory and the real-time motion trajectory are determined by the head posture sequence data, so as to detect whether the gaze object of the user changes, and then the real-time audio-video data of the new target gaze object is reacquired in time for processing when it is detected that the gaze object of the user changes. Through the above-mentioned manner, the application can actively perceive the intention of the user, and acquire the real-time audio-video data of the switched target speaking object in time for speech recognition, and then the speech source can be switched in real time following the change of the object of attention of the user.
[0088] Here, the first head posture sequence data can be used to start timing when determining the target gaze object, collect head posture data in a 1.5s time window, and form the head posture sequence data, denoted as the first head posture sequence data. The first head posture sequence data includes multiple frames of head posture data, wherein the head posture data can be a quaternion.
[0089] As an implementation, the IMU can be used to sample at a sampling rate of 1000Hz.
[0090] The determination process of the real-time motion trajectory is as follows: each frame of head posture data in the first head posture sequence data is vectorized to obtain a first vector sequence as the real-time motion trajectory.
[0091] The vectorization of each frame of head posture data is as follows: the forward unit axis (0, 0, 1) in the XR device's own coordinate system is regarded as the "XR device front" basis vector, and the unit orientation vector is calculated using the quaternion rotation formula, denoted as , wherein the quaternion rotation formula is .
[0092] Step S160, real-time acquisition of the second head posture sequence data of the user, and detection of whether the gaze object of the user changes according to the second head posture sequence data and the current reference trajectory.
[0093] Real-time acquisition of the second head posture sequence data of the user, determination of the real-time motion trajectory according to the second head posture sequence data; encoding the real-time motion trajectory into a real-time feature vector, and encoding the current reference trajectory into a current reference feature vector; calculating the deviation degree according to the real-time feature vector and the current reference feature vector; detecting whether the deviation degree is greater than a preset deviation threshold to detect whether the gaze object of the user changes. The specific execution process can refer to the following embodiments, which are not described here.
[0094] If it is detected that the gaze object of the user does not change, step S130 of acquiring real-time audio data and real-time video data of the target gaze object is continued.
[0095] If it is detected that the gaze object of the user changes, step S110 of acquiring the current scene image and the current head posture data of the user is re-executed.
[0096] If it is detected that the gaze object of the user does not change, the real-time audio and video data of the current target gaze object is continuously acquired, and then recognition and output are performed. If it is detected that the gaze object of the user changes, the target gaze object is re-determined, specifically, the current scene image and the current head posture data of the user are re-acquired to determine a new target gaze object, and then subsequent steps are performed.
[0097] In this embodiment, the first head posture sequence data and the second head posture sequence data of the user are acquired to detect whether the gaze object of the user changes by using the head posture change. The user does not need any additional operation (such as manual selection, voice instruction), as long as the user naturally turns the head to look at another speaking object, the application can automatically identify this intention and quickly switch the target speaking object, solve the problem of inconvenient interaction and lack of automation in the prior art, and improve the naturalness of human-computer interaction.
[0098] In an embodiment, the step S160 comprises a step S161, a step S162, a step S163 and a step S164.
[0099] In step S161, the second head posture sequence data of the user is acquired in real time, and a real-time motion trajectory is determined according to the second head posture sequence data.
[0100] Here, the head posture data of the user is acquired in real time, and the head posture sequence data composed of the head posture data in a preset time window is recorded as the second head posture sequence data, which is used to monitor whether the user switches the target gaze object. The second head posture sequence data comprises a plurality of frames of head posture data, wherein the head posture data can be a quaternion.
[0101] As an implementation manner, the IMU can be used to sample at a sampling rate of 1000Hz.
[0102] The determination process of the real-time motion trajectory is as follows: each frame of head posture data in the second head posture sequence data is vectorized to obtain a second vector sequence as the real-time motion trajectory.
[0103] The vectorization of each frame of head posture data is as follows: the forward unit axis (0, 0, 1) in the XR device coordinate system is regarded as the "XR device front" basis vector, and the unit orientation vector is calculated using the quaternion rotation formula, denoted as , wherein the quaternion rotation formula is .
[0104] In step S162, the real-time motion trajectory is encoded into a real-time feature vector, and the current reference trajectory is encoded into a current reference feature vector.
[0105] Since the real-time motion trajectory is the second vector sequence, it can be encoded into a fixed-length feature vector by an LSTM-Encoder (Long Short-Term Memory-Encoder), denoted as a real-time feature vector γ(t). The LSTM-Encoder is a neural network module composed of LSTM units, and its task is to understand the context information of the entire input sequence and encode it into a meaningful fixed-size feature vector.
[0106] Specifically, the real-time feature vector γ(t) is a 32-dimensional feature vector, including four types of statistical features: spatial features, motion features, geometric features, and time sequence features. Among them, the spatial features, i.e., the start and end vectors, are 6-dimensional; the motion features include the mean and variance of the angular velocity, which are 4-dimensional, and the peak angular velocity, which is 2-dimensional; the geometric features, i.e., the segmented curvature, are 8-dimensional; and the time sequence features, i.e., the LSTM hidden state, are 12-dimensional. Through fixed-length multi-dimensional features, the head motion can be better described. When using a 32-dimensional feature vector, through testing, the delay can be ≤50 ms, and the head behavior prediction accuracy can be ≥95%.
[0107] Further, before encoding by the LSTM-Encoder, the second vector sequence is sequentially subjected to filtering and compression processing to obtain a processed second vector sequence.
[0108] As an implementation, the filtering processing can use a sliding average filter or a first-order Kalman filter. Through filtering processing, high-frequency jitter and noise in the second vector sequence can be suppressed, making the data smoother and improving the data quality, preparing for subsequent compression.
[0109] As an implementation, the compression processing can use the RDP (Ramer-Douglas-Peucker) algorithm. A threshold ε=0.5° can be set to retain key inflection points and compress redundant data. Through compression processing, the data volume can be greatly reduced, and non-information points can be removed, so that the number of data points that the subsequent LSTM-Encoder needs to process is greatly reduced, and the computational burden and time of training and inference can be significantly reduced.
[0110] Similarly, the current reference trajectory is encoded into a current reference feature vector, denoted as γ0.
[0111] Step S163, the deviation degree is calculated according to the real-time feature vector and the current reference feature vector.
[0112] The calculation formula of the deviation degree is as follows:
[0113] .
[0114] wherein, is the deviation degree; w1, w2 and w3 are preset weight values, w1=0.6, w2=0.25, w3=0.15; K t is the real-time curvature, i.e., the curvature in the real-time feature vector; K0 is the reference curvature, i.e., the curvature in the current reference feature vector; is the real-time peak angular velocity, i.e., the peak angular velocity in the real-time feature vector; is the reference peak angular velocity, i.e., the peak angular velocity in the current reference feature vector; C s is the cosine similarity, which can be calculated by the following formula:
[0115] .
[0116] Further, after the current reference feature vector is calculated, a historical reference feature vector can be obtained, the current reference feature vector and the historical reference feature vector are fused to obtain a fused reference feature vector, and then the deviation degree is calculated according to the real-time feature vector and the fused reference feature vector, so that the trajectory model continuously fits the real habits of the user, thereby making the switching detection result of the gaze object more accurate.
[0117] Specifically, when fusing, EMA (Exponential Moving Average) fusion can be performed. The specific formula is as follows:
[0118] .
[0119] wherein, denotes the fused reference feature vector, denotes the current reference feature vector, denotes the historical reference feature vector, and a is a preset weight.
[0120] The value of a can be adaptively adjusted, and the adjustable range is 0.1-0.3. The adjustment strategy is: 1) for the first use, a=0.3, to quickly establish an initial template; 2) for 3 consecutive similar trajectories (i.e., cosine similarity>90%), a=max(a 0.8, 0.1), to reduce the learning rate and stabilize the mature template; 3) for trajectory mutation (i.e., cosine similarity<85%), a=min(0.3, a*1.5), to sensitively adapt to new habits; 4) for a high-noise environment (i.e., sigma h >10°), freeze the update (a=0) to prevent vibration from polluting the template.
[0121] In step S164, it is detected whether the deviation degree is greater than a preset deviation threshold, to detect whether the gaze object of the user changes.
[0122] After the deviation degree is calculated, it is detected whether the deviation degree is greater than a preset deviation threshold, so as to detect whether the gaze object of the user changes. If the deviation degree is greater than the preset deviation threshold, it is determined that the gaze object of the user changes. If the deviation degree is less than or equal to the preset deviation threshold, it is determined that the gaze object of the user does not change.
[0123] The preset deviation threshold can be a fixed value, and can also be a dynamic threshold, so as to take into account stability and jitter tolerance. Specifically, the preset deviation threshold can be determined according to the statistical head micro-motion standard deviation (denoted as σ) of the user, and the formula is as follows: wherein σ is the preset deviation threshold, σ0 is a basic deviation threshold, which is a preset value, and σ is the head micro-motion standard deviation.
[0124] In the embodiment, the trajectory deviation threshold strategy is introduced in the target switching mechanism, which can sensitively identify the transfer of the user's gaze direction, and then switch the target gaze object.
[0125] In an embodiment, the step S164 includes a step S1641, a step S1642, a step S1643 and a step S1644.
[0126] The step S1641 detects whether the deviation degree is greater than a preset deviation threshold.
[0127] The step S1642 detects that the deviation degree is less than or equal to the preset deviation threshold, and determines that the change of the gaze object of the user is not detected.
[0128] The step S1643 detects that the deviation degree is greater than the preset deviation threshold, and starts timing to monitor whether the deviation degree in a preset time is greater than the preset deviation threshold.
[0129] The step S1644 detects that the deviation degree in the preset time is greater than the preset deviation threshold, and determines that the change of the gaze object of the user is detected.
[0130] In the embodiment, considering that the user may quickly scan the environment or briefly look at other objects in the process of obtaining the voice of the target gaze object, in order to avoid frequent switching, the gaze confirmation strategy is added on the basis of the above trajectory deviation threshold strategy, which can sensitively and stably identify the transfer of the user's attention.
[0131] Specifically, it is detected whether the deviation degree is greater than the preset deviation degree threshold value. If it is detected that the deviation degree is less than or equal to the preset deviation degree threshold value, it is determined that the change of the user's gaze object is not detected. If it is detected that the deviation degree is greater than the preset deviation degree threshold value, the timing is started, and it is monitored whether the deviation degrees in the preset time are all greater than the preset deviation degree threshold value. If it is monitored that the deviation degrees in the preset time are all greater than the preset deviation degree threshold value, it is determined that the change of the user's gaze object is detected. If it is monitored that the deviation degrees in the preset time exist less than or equal to the preset deviation degree threshold value, it is determined that the change of the user's gaze object is not detected. The preset time can be set according to user habits, and for example, can be set to 300 milliseconds.
[0132] In the target switching mechanism in this embodiment, the gaze confirmation strategy is added on the basis of the trajectory deviation threshold strategy, which can avoid frequent mis-switching, so that the transfer of user attention can be sensitively and stably recognized. Through the above-mentioned manner, balance can be achieved between response speed and misjudgment avoidance, which can respond to user intention in milliseconds and will not shake back and forth or switch when quickly scanning the environment due to slight head movement.
[0133] Based on any of the above embodiments, referring to Figure 4 , Figure 4 is a third flowchart of the speech processing method provided by the present application, and step S140 includes step S141, step S142 and step S143.
[0134] In step S141, the real-time audio data is processed by a time sliding window, and the real-time video data is processed by a frame.
[0135] Considering the delay problem, in this embodiment, when the target audio signal of the target gaze object is determined by using the speech separation model, a streaming processing mode is adopted. For example, if the recording time of the audio and video data is set to 3s, in the prior art, the 3s audio and video data recorded is generally input into the speech separation model for processing. In this embodiment, the streaming processing mode is adopted, and every time 400ms of real-time audio and video data (including real-time audio data and real-time video data) recorded is input into the speech separation model for processing during the recording process of the audio and video data. Through streaming processing, the audio and video data can be processed immediately after being generated in the early stage of acquisition, so that fast processing of the audio and video data can be realized, the delay is greatly reduced, and the user experience is improved.
[0136] In the streaming processing, the input of the speech separation model is generally processed in the form of a chunk (block) in a time window. Therefore, after the real-time audio data is acquired, the real-time audio data is processed by blocking, and the real-time video data acquired synchronously is processed by framing.
[0137] Further, considering the latency problem and user experience, the audio length of the chunk processing can be determined based on the key indicator of streaming efficiency, i.e., the end-to-end latency.
[0138] Specifically, the end-to-end latency consists of the following: ; wherein, is the end-to-end latency, is the time of each chunk audio, is the model forward calculation time. Considering the user experience problem, should be controlled to be < 500 ms, should be controlled to be < 100 ms, corresponding to can be set to 400 ms.
[0139] Further, the overlap ratio (overlap) is set to 50%, i.e., the current chunk audio and the previous chunk audio overlap by half. By adopting the 50% overlap sliding window mechanism, the adjacent two times of processing have a half overlap, so that a part of the context can be reserved between each operation, which helps to smooth the transition and reduce the output delay.
[0140] Exemplarily, the commonly used audio sampling rate is 16 kHz; the length of each chunk audio is 6400, i.e., the audio length of the chunk audio processed each time is 400 ms; the sliding step (hop) is considered to be 3200, and each time the audio is slid back by 3200 sampling points (i.e., 200 ms); the overlap ratio (overlap) is 50%, i.e., the current chunk audio and the previous chunk audio overlap by half. Correspondingly, the frame processing data of real-time video data is as follows: assuming that the video frame rate is 25 FPS, corresponding to one frame every 40 ms, then each 400 ms chunk audio contains about 10 image frames. When processing the nth chunk audio, the real-time video data corresponding to the time period is synchronously collected, and the video frame sequence is obtained by frame division, and then input into the model.
[0141] Further, for the first frame, we can use the "half-padding" method to occupy the position of the latter half of the samples that have not yet arrived with "useless" or zero signals. In this way, when the real 200 ms real-time audio data arrives, a 400 ms chunk audio can be immediately formed and input into the speech separation model, so as to further reduce the end-to-end latency to < 300 ms.
[0142] Further, a sliding window (overlap) can also be used to further reduce the latency. Specifically, the audio buffer area is constantly updated with a sliding step. That is, in the nth processing period, the audio data from the time point to audio data, constituting the sub-chunk audio of the sub-process . Wherein, is the sliding step length, is the audio block length (sample number) of each processing.
[0143] Step S142, input the real-time audio data processed by the sub-chunk processing and the real-time video data processed by the frame processing into the speech separation model, to obtain the sub-chunk audio signal corresponding to the target gaze object output by the speech separation model.
[0144] Then, the real-time audio data processed by the sub-chunk processing and the real-time video data processed by the frame processing are input into the speech separation model, and the sub-chunk audio signal corresponding to the face information is determined by using the speech separation model.
[0145] Specifically, the speech separation model includes an audio encoder, a visual encoder, an audio-video fusion module, a mask estimation network, and an audio decoder. First, the real-time audio data is input into the audio encoder to obtain the audio features output by the audio encoder, and at the same time, the real-time video data is input into the visual encoder to obtain the face visual features output by the visual encoder. Then, the audio features and the face visual features are input into the audio-video fusion module to obtain the audio-video fusion features output by the audio-video fusion module. Then, the audio-video fusion features are input into the mask estimation network to obtain the time-frequency mask output by the mask estimation network. Finally, the time-frequency mask is input into the audio decoder to obtain the target gaze object corresponding to the sub-chunk audio signal output by the audio decoder. The specific execution process can be referred to the following embodiments, which will not be repeated here.
[0146] Step S143, the sub-chunk audio signals are spliced by a smoothing window to obtain the target audio signal of the target gaze object.
[0147] The time-domain output signal (i.e. sub-chunk audio signal) generated for each sub-chunk audio is denoted as , and a smoothing window (such as a Hanning window) is applied for processing: ; wherein, is the smoothed signal after applying a window function to the in-block signal , and is the value of the smoothing window function at time point t.
[0148] The is added and spliced with the overlapping section of the previous chunk to ensure seamless connection. Specifically as follows:
[0149] ;
[0150] Wherein, output represents the processing result, n represents the nth processing period, is the smoothed signal after applying a window function to the in-block signal the smoothed signal after applying the window function, for a sliding step.
[0151] In this embodiment, through stream processing, the audio and video data can be quickly processed, the delay is greatly reduced, and the user experience is improved.
[0152] In an embodiment, the speech separation model comprises an audio encoder, a visual encoder, an audio-video fusion module, a mask estimation network, and an audio decoder, and the step S142 comprises steps S1421, S1422, S1423, S1424, and S1425.
[0153] In step S1421, the real-time audio data processed by the block processing is input into the audio encoder to obtain the audio features output by the audio encoder.
[0154] Here, the audio features include but are not limited to short-term audio features, long-distance dependency features, and audio features fused based on short-term audio features and long-distance dependency features.
[0155] As an implementation, when the audio features are short-term audio features, the audio encoder can be a U-Net (U-shaped network) encoder to extract short-term audio features through the U-Net encoder.
[0156] As another implementation, when the audio features are long-distance dependency features, the audio encoder can be a time-frequency transformer network to extract long-distance dependency features through the time-frequency transformer network.
[0157] As yet another implementation, when the audio features are audio features fused based on short-term audio features and long-distance dependency features, the audio encoder comprises a U-Net encoder, a time-frequency transformer network, and an audio fusion module.
[0158] The U-Net is a convolutional neural network architecture mainly used for separation, and the U-Net adopts an encoder-decoder structure. The real-time audio data processed by the block processing can be input into the U-Net encoder to obtain the short-term audio features output by the U-Net encoder. Specifically, the audio waveform of the input real-time audio data processed by the block processing can be obtained through 1D convolution and residual blocks by the U-Net encoder to obtain short-term audio features. The short-term audio features mainly include time domain features and frequency domain features, and these features focus on the change characteristics of the real-time audio data in time and frequency.
[0159] The time-frequency transformer network is used for transforming the audio waveform of the real-time audio data into a time-frequency spectrum, and long-distance dependency features can be extracted by the transformer. Specifically, the block-processed real-time audio data is input into the time-frequency transformer network, and long-distance dependency features output by the time-frequency transformer network are obtained. The long-distance dependency features can reflect the complex correlations across time or frequency in the block-processed real-time audio data.
[0160] The audio fusion module is used for feature fusion of the short-term audio features and the long-distance dependency features. The short-term audio features and the long-distance dependency features are input into the audio fusion module, and audio features output by the audio fusion module are obtained. Specifically, the audio fusion module can sequentially perform linear projection and splicing processing on the short-term audio features and the long-distance dependency features to obtain the audio features. Specifically, ; wherein, is the audio feature, is the short-term audio feature, is the long-distance dependency feature, is a linear projection function of the short-term audio feature, is a linear projection function of the long-distance dependency feature, and Concat is a connection function.
[0161] In this embodiment, by extracting and fusing the short-term audio features and the long-distance dependency features of the block-processed real-time audio data, both local feature details and global feature information are considered, which can help improve the accuracy of the speech separation model recognition result.
[0162] In step S1422, the frame-processed real-time video data is input into the visual encoder, and face visual features output by the visual encoder are obtained.
[0163] Here, the face visual features include but are not limited to: face appearance features, lip movement sequence features, and face visual features fused based on the face appearance features and the lip movement sequence features.
[0164] As an implementation manner, when the face visual features are based on the face appearance features, the visual encoder can be a ResNet-18 network structure (a kind of deep residual network), so as to extract the face appearance features through the network structure.
[0165] As a further implementation manner, when the facial visual feature is a lip movement sequence feature, the visual encoder can be a network structure composed of a 3D-Conv (3D-Convolutional Neural Networks), a ShuffleNet V2 (an efficient convolutional neural network), and a TCN (Temporal Convolutional Network), so as to extract the lip movement sequence feature through the network structure.
[0166] As another implementation manner, when the facial visual feature is a facial visual feature fused based on a facial appearance feature and a lip movement sequence feature, the visual encoder can include a facial attribute subnetwork, a lip movement subnetwork, and a visual fusion module.
[0167] The facial attribute subnetwork can be a ResNet-18 network structure (a deep residual network) for extracting the facial appearance feature. The real-time video data processed by the frame division can be input into the facial attribute subnetwork to obtain the facial appearance feature output by the facial attribute subnetwork.
[0168] The lip movement subnetwork can be a network structure composed of a 3D-Conv (3D-Convolutional Neural Networks), a ShuffleNet V2 (an efficient convolutional neural network), and a TCN (Temporal Convolutional Network) for extracting the lip movement sequence feature. The real-time video data processed by the frame division can be input into the lip movement subnetwork to obtain the lip movement sequence feature output by the lip movement subnetwork.
[0169] The visual fusion module can be a transformer network or a small attention network for evaluating the visual frame quality and outputting the facial visual feature. The facial appearance feature and the lip movement sequence feature can be input into the visual fusion module to obtain the facial visual feature output by the visual fusion module. Specifically, the facial appearance feature and the lip movement sequence feature are weighted and fused to obtain the facial visual feature. Specifically as follows:
[0170] ;
[0171] wherein, is the facial visual feature, is the facial appearance feature, is a weight coefficient of the facial appearance feature, is the lip movement sequence feature, is a weight coefficient of the lip movement sequence feature.
[0172] In this embodiment, the face appearance feature and the lip movement sequence feature of the real-time video data after the frame processing are extracted and fused to obtain the face visual feature. The fusion of the multiple face features related to the speech separation can help improve the accuracy of the speech separation model recognition result.
[0173] In step S1423, the audio feature and the face visual feature are input into the audio-video fusion module to obtain the audio-video fusion feature output by the audio-video fusion module.
[0174] Specifically, the audio feature and the face visual feature are first input into the audio-video fusion module, and the audio feature and the face visual feature are dynamically fused by the audio-video fusion module to obtain the current audio-video fusion feature.
[0175] The dynamic cross-modal gating (DCG) module can be included in the audio-video fusion module, which is used to determine the cross-modal fusion weight based on the reliability of the audio feature and the reliability of the face visual feature obtained by dynamic evaluation. Specifically as follows:
[0176] ;
[0177] Wherein, is the cross-modal fusion weight, is the reliability of the audio feature, is the reliability of the face visual feature, and σ is the Sigmoid function, , and b are learnable parameters. Wherein, is used to measure the quality of the audio signal or the reliability of the audio branch output at the current time, generally taking a value between [0, 1]. A small MLP (Multi-Layer Perceptron) is connected at the end of the U-Net / transformer audio branch, and the feature vector is input into it. After two fully connected layers and Sigmoid activation function, a scalar is directly predicted, which is . is used to measure the reliability of the output of the visual branch (such as face / lip movement) at the current time, also between [0, 1]. Similarly, a small MLP can be added at the end of the visual feature branch to obtain .
[0178] Then, based on the cross-modal fusion weight, the audio feature and the face visual feature are dynamically fused to obtain the current audio-video fusion feature. Specifically as follows:
[0179] ;
[0180] wherein, is the current audio-video fusion feature, is the audio feature, is the facial visual feature. represents that the audio feature and the facial visual feature are spliced, and then the result obtained by learning fusion is obtained by using several layers of MLP. represents that the audio feature is linearly projected, and the vector dimension is kept unchanged or mapped to the target dimension.
[0181] Then, the historical audio-video fusion features are weighted and fused from the multi-level feature memory library based on the current audio-video fusion feature to obtain an audio-video fusion feature; wherein the historical audio-video fusion features include one or more of short-term memory audio-video fusion features, long-term memory audio-video fusion features, and global memory audio-video fusion features.
[0182] Here, the multi-level feature memory library includes one or more of short-term memory audio-video fusion features, long-term memory audio-video fusion features, and global memory audio-video fusion features, so the weighted and fused historical audio-video fusion features can also be one or more of them.
[0183] Among them, the short-term memory audio-video fusion feature can be an audio-video fusion feature generated in real time in the recent few windows, the long-term memory audio-video fusion feature can be an audio-video fusion feature of the target object accumulated in the current session, and the global memory audio-video fusion feature can be an audio-video fusion feature of the target object across multiple sessions.
[0184] The audio-video fusion module can also include a dynamic window fusion (DWF) module for cross-window weighted fusion processing of the historical audio-video fusion features and the current audio-video fusion features. Specifically, the historical window features can be dynamically integrated based on a self-attention mechanism, and high-quality information in the historical multiple windows can be adaptively selected, and cross-window fusion is realized through a transformer. The specific weighted fusion process is as follows:
[0185] ;
[0186] wherein, is the audio-video fusion feature, is the current audio-video fusion feature, is the short-term memory audio-video fusion feature, is the long-term memory audio-video fusion feature, is the global memory audio-video fusion feature.
[0187] By performing cross-window weighted fusion processing on the current audio-video fusion feature and the historical audio-video fusion feature, a high-quality audio-video fusion feature can be fused, thereby further improving the accuracy of the speech separation model recognition result.
[0188] In step S1424, the audio-video fusion feature is input into the mask estimation network to obtain a time-frequency mask output by the mask estimation network.
[0189] The mask estimation network (Mask Estimator) is used for automatic estimation and generation of a time-frequency mask of a target object.
[0190] As an implementation, in a non-streaming processing case, the time-frequency mask is determined according to the following manner: . Wherein, is the time-frequency mask, is the audio feature, is an audio-video fusion feature obtained by cross-modal dynamic gating fusion technology, is a frequency domain mask function.
[0191] As an implementation, in a streaming processing case, the time-frequency mask is determined according to the following manner: . Wherein, is the time-frequency mask, is the audio feature, is an audio-video fusion feature obtained by weighted fusion processing, is a frequency domain mask function.
[0192] In step S1425, the time-frequency mask is input into the audio decoder to obtain a block audio signal corresponding to the target gaze object output by the audio decoder.
[0193] The audio decoder is used to restore the time-frequency mask to a time-domain waveform, and then output the corresponding audio signal. The process of restoring the time-frequency mask to the time-domain waveform is as follows:
[0194] s ;
[0195] Wherein, s is the time-domain waveform, is the time-frequency mask, is a short-time Fourier transform of the block-processed real-time audio data, is a Hadamard product, is an inverse short-time Fourier transform function. That is, after multiplying the time-frequency mask and the STFT of the block-processed real-time audio data, the time-domain signal is reconstructed by inverse STFT.
[0196] In this embodiment, the real-time audio and video data is processed by the voice separation model to accurately determine the audio signal corresponding to the target gaze object. The optimized model structure and algorithm of this embodiment meet the embedded operation requirements, and the actual running power consumption is low, and the whole machine inference power consumption is less than 400mW. This voice separation model with strong real-time performance and low power consumption can run for a long time on a battery-powered XR device without excessive power consumption or significant heat generation.
[0197] In an embodiment, the voice separation model is trained based on sample mixed audio, sample face, and sample face audio signal annotation results.
[0198] As an implementation, the sample mixed audio is audio mixed with multiple human voices and noise.
[0199] In an embodiment, the voice separation model is trained using a preset loss function, which is determined based on a voice separation loss, a speaker recognition loss, and a visual quality evaluation loss.
[0200] In a specific embodiment, the preset loss function is determined based on the following formula:
[0201] ;
[0202] wherein, is the total loss, is the voice separation loss, is the speaker recognition loss, is the visual quality evaluation loss, and α and β are two hyperparameters for balancing the weights of each sub-loss in the overall optimization process. They need to be selected through small-scale search (such as grid search or Bayesian optimization) before training, so that the final model achieves the optimal balance between voice separation quality, speaker discrimination ability, and visual fusion rationality.
[0203] wherein, the voice separation loss is determined based on the following formula:
[0204] ;
[0205] wherein, is the estimated signal, and s is the reference signal (i.e., the clean signal); is the target component, which refers to the part that aligns with the reference signal as much as possible in the reference signal direction, and measures how much of the estimated signal is truly aligned to the reference signal; is the error component, which refers to the part of the estimated signal that is not aligned with the reference signal, and includes residual noise, distortion, or content other than the source signal. is determined based on the following formula: .
[0206] The speaker recognition loss is determined based on the following formula:
[0207] ;
[0208] wherein C is the total number of speaker categories, is an indicator variable of the true label on the c-th category, usually encoded by one-hot, if the sample belongs to the speaker category c, = 1, otherwise = 0. is the predicted probability of the speech separation model that the sample belongs to the c-th category, usually normalized by Softmax.
[0209] The visual quality evaluation loss is determined based on the following formula:
[0210] ;
[0211] wherein is the true visual reliability, is the predicted visual reliability.
[0212] The above preset loss function integrates the speech separation loss , the speaker recognition loss and the visual quality evaluation loss , which can better supervise the model optimization processing result, thereby further improving the processing effect of the speech separation model.
[0213] Based on any of the above embodiments, step S120 includes steps S121, S122, S123, S124, S125 and S126. It should be noted that the execution order of steps S121-S124 and S125 is not sequential.
[0214] Step S121, face detection is performed on the current scene image to obtain candidate objects and their pixel coordinates.
[0215] When face detection is performed, a face detection algorithm such as MTCNN (Multi-task Convolutional Neural Network), face recognition network, etc. can be used to locate all face regions in the current scene image, output two-dimensional bounding boxes, and the faces in each two-dimensional bounding box are candidate objects. The pixel coordinates of the candidate objects are the center point coordinates in the two-dimensional bounding box, denoted as (u, v).
[0216] Step S122, a coordinate system is established according to the X axis and the Z axis, with the user's head as the origin, the front of the user's head as the X axis, and the vertical upward direction of the origin as the Z axis.
[0217] Step S123, the relative coordinates of the candidate object in the coordinate system are determined according to the pixel coordinates and the current scene image.
[0218] Then, a coordinate system, i.e., a device coordinate system, is established according to the X axis and the Z axis, with the user's head as the origin, the front of the user's head as the X axis, and the vertical upward direction of the origin as the Z axis. Further, the relative coordinates of the candidate object in the coordinate system are obtained according to the pixel coordinates and the current scene image, denoted as P_face.
[0219] Specifically, the depth value of the candidate object is obtained according to the current scene image, and the three-dimensional coordinates of the candidate object in the camera coordinate system are obtained according to the pixel coordinates and the depth value, denoted as P_camera. Then, the three-dimensional coordinates P_camera in the camera coordinate system are transformed to the device coordinate system to obtain the relative coordinates of the candidate object in the device coordinate system.
[0220] As an implementation manner, if the current scene image is obtained by an RGB camera, the current scene image is a two-dimensional image. At this time, a monocular depth estimation model can be used to generate a pixel-level depth map according to the current scene image, and then the depth value is obtained according to the pixel-level depth map. Alternatively, the face width is calculated according to the two-dimensional bounding box in the current scene image, and then the face width is substituted into a preset depth value calculation formula to determine the depth value. The preset depth value calculation formula is: z=k1·W real1 / W px1 . Wherein, z is the depth value, k1 is a preset calibration coefficient, W real1 is a preset real face width, and W px1 is the calculated face width. Alternatively, the interpupillary distance is obtained according to the two-dimensional bounding box in the current scene image, and then the interpupillary distance is substituted into a preset depth value calculation formula to determine the depth value. The preset depth value calculation formula is: z=k2·W real2 / W px2 . Wherein, z is the depth value, k2 is a preset calibration coefficient, W real2 is a preset interpupillary distance, and W px2 is the obtained interpupillary distance.
[0221] As another implementation manner, if the current scene image is obtained by an RGB-D camera, the current scene image is a depth map. At this time, the depth value of the center of the two-dimensional bounding box marked in the current scene image can be directly read, denoted as z.
[0222] As another implementation, if the current scene image is obtained by a binocular camera, and the current scene image comprises a first scene left image and a first scene right image, a horizontal pixel coordinate u_left of the candidate object in the first scene left image is obtained, and a horizontal pixel coordinate u_right of the candidate object in the second scene right image is obtained, subtraction is performed on u_left and u_right to obtain a disparity d, i.e., d = u_left - u_right. Then, according to the preset focal length, the preset baseline distance and the disparity, a depth value is calculated. The preset focal length and the preset baseline distance are respectively a focal length and a baseline distance corresponding to the binocular camera, and the depth value z = (f·b) / d, f is the preset focal length, and b is the preset baseline distance.
[0223] The calculation process of the three-dimensional coordinates P_camera of the candidate object in the camera coordinate system is as follows:
[0224] The pixel coordinates (u, v) are converted into three-dimensional coordinates P_camera in the camera coordinate system in combination with a camera intrinsic matrix (denoted as K). The specific formula is as follows:
[0225] .
[0226] Finally, according to a camera-XR device extrinsic matrix (denoted as T), P_camera is transformed into a device coordinate system to obtain a relative coordinate P_face of the candidate object in the device coordinate system. The specific formula is as follows: P_face = T·P_camera.
[0227] In step S124, a face direction vector of the candidate object is calculated according to the relative coordinate and a reference coordinate of the head of the user.
[0228] The relative coordinate P_face is subtracted from the reference coordinate of the head of the user (denoted as C_head) and is unitized to obtain a face direction vector (denoted as vec{d0}) of the candidate object. The specific formula is as follows:
[0229] vec{d0} = (P_face - C_head) / |P_face - C_head|.
[0230] In step S125, a head orientation vector of the user is calculated according to the current head posture data.
[0231] The current head posture data is a quaternion (denoted as q) representing a head posture output by an IMU. The quaternion q is obtained by fusion of a gyroscope and an accelerometer.
[0232] As an implementation, the forward unit axis (0, 0, 1) in the XR device's own coordinate system is regarded as the "XR device front" basis vector, and the head orientation vector of the user is calculated using the quaternion rotation formula , wherein the quaternion rotation formula is .
[0233] As another implementation, the equivalent direction cosine matrix multiplication can also be used to calculate the head orientation vector of the user according to the current head pose data.
[0234] It should be noted that the physical meaning of the head orientation vector of the user is: if the user maintains the current head posture, the infinite point in the direction of is the center of the user's line of sight.
[0235] Step S126, calculate the included angle between the head orientation vector and the face direction vector, and determine the target gaze object from the candidate objects according to the calculated included angle.
[0236] After obtaining the head orientation vector of the user and the face direction vector of the candidate object, the included angle between the two vectors is calculated. The specific calculation formula is: α = arccos (vec{d0}·vec{h}); wherein, α is the included angle, and arccos is the inverse sine function.
[0237] As an implementation, when determining the target gaze object, if the candidate objects include multiple, that is, the included angles include multiple, the candidate object corresponding to the smallest included angle can be selected as the target gaze object. If the candidate object is only one, the candidate object can be directly selected as the target gaze object. It should be understood that after step S121, the number of candidate objects can be detected, and if it is one, it can be directly determined as the target gaze object.
[0238] As another implementation, when determining the target gaze object, the smallest included angle is selected from the calculated included angles, denoted as α min . Further, it is judged whether the smallest included angle is less than or equal to a preset threshold value to determine whether the head line of sight of the user has been aligned to the target object, wherein the preset threshold value (denoted as α_align) can be a fixed value, for example, it can be set to any value in 6°~8°, and it can also be determined according to the statistical standard deviation (denoted as σ h ) of the head micro-motion of the user, and specifically it can be set to α_align = max (4°, 2σ h), to balance stability and jitter tolerance. If the minimum included angle is less than or equal to a preset threshold, the candidate object corresponding to the minimum included angle is determined as the target gaze object; if the minimum included angle is greater than the preset threshold, it is determined that there is no target gaze object in the current scene image. At this time, a prompt information can be generated to make the user gaze at the target object, and then further acquire the corresponding scene image to determine the target gaze object.
[0239] In the embodiment, the candidate object and the face direction vector thereof are acquired according to the current scene image and the current head pose data, and at the same time, the head orientation vector of the user is acquired according to the current head pose data, and then the included angle between the head orientation vector and the face direction vector is calculated to determine the target gaze object from the candidate objects according to the calculated included angle. Through the above-mentioned manner, the target gaze object can be accurately determined from the current scene image, especially in a multi-person scene.
[0240] Based on any of the above embodiments, after step S140, the voice processing method further comprises step S170.
[0241] Step S170, intensity change processing is performed on the target audio signal of the target gaze object in the real-time audio data.
[0242] In the embodiment, considering that the background noise and / or the speaking sound of the non-target gaze object has a large interference, after the target audio signal of the target gaze object is identified, the intensity change processing is further performed on the target audio signal of the target gaze object in the real-time audio data. The intensity change processing manner can include: 1) only retaining the target audio signal of the target gaze object, eliminating the audio signals of other objects and the background noise; 2) emphasizing the target audio signal of the target gaze object, and weakly retaining the audio signals of other objects and the background noise.
[0243] Through the intensity change processing, the volume of the non-target gaze object and the background noise can be reduced or even eliminated in the audio output based on the target audio signal after the intensity change processing, so as to highlight the volume of the target gaze object, facilitate the user to better hear the sound of the target gaze object, reduce the interference of other interference sounds, and thus improve the user's use experience. In actual application, for example, in business meeting, social gathering or education training scene, the user wearing the XR device (such as XR glasses) can focus on the sound of the target gaze object, even if there are many people talking around. This will significantly improve the efficiency and quality of communication among many people, and reduce the situation of "misunderstanding" or "missing". At the same time, for noisy environment (such as construction site, banquet, etc.), the present application can also serve as a directional noise reduction listening device to improve the user's perception ability of specific sound sources.
[0244] In a specific embodiment, the real-time audio data comprises audio signals of at least two objects, and the intensity change processing of the target audio signal of the target gaze object is implemented by setting weight values for the audio signals of the at least two objects.
[0245] Further, considering that in a multi-person scenario, the audio to be processed can comprise audio signals of multiple objects, the intensity change processing of the target audio signal of the target gaze object can be implemented by setting weight values for the audio signals of the at least two objects.
[0246] It can be understood that if the real-time audio data only comprises an audio signal of one object, i.e., the target audio signal of the target gaze object, no processing is needed. If the real-time audio data only comprises an audio signal of one object, but also comprises background noise, the background noise can be removed or weakly reserved, and similarly, the intensity change processing can also be implemented by setting weight values. Specifically, different weight values can be set for the target audio signal of the target gaze object and the audio signal of the background noise.
[0247] In a specific embodiment, the step S170 can comprise:
[0248] When the weight value of the audio signal of the non-target gaze object is set to 0, the target audio signal of the target gaze object is separated from the real-time audio signal corresponding to the real-time audio data; or,
[0249] When the weight value of the audio signal of the target gaze object is set to be greater than the weight value of the audio signal of the non-target gaze object, the target audio signal of the target gaze object is enhanced in the real-time audio signal corresponding to the real-time audio data; wherein the non-target gaze object is an object other than the target gaze object among the at least two objects.
[0250] It should be noted that in this embodiment, the influence of background noise is not considered, and only the scenario in which multiple objects speak is considered.
[0251] As an implementation, the weight value of the audio signal of the non-target gaze object can be set to 0 to eliminate the interference sound of other objects. The weight value of the audio signal of the target gaze object can be set to 1. In this case, the target audio signal of the target gaze object can be separated from the real-time audio signal corresponding to the real-time audio data.
[0252] As an implementation, the weight value of the audio signal of the target gaze object can be set to be greater than the weight value of the audio signal of the non-target gaze object, so as to emphasize the voice of the target gaze object and weakly reserve the voice of the non-target gaze object. For example, the weight value of the audio signal of the target gaze object can be set to 1, and the weight value of the audio signal of the non-target gaze object is in a range greater than 0 and less than 1. In this case, in the real-time audio signal corresponding to the real-time audio data, the target audio signal of the target gaze object is enhanced. Specifically, the sampling points of the target audio signal and the non-target audio signal are weighted and summed to obtain an output audio signal, wherein the non-target audio signal is the audio signal of the object other than the target gaze object in the real-time audio data.
[0253] Further, after the intensity change processing is performed, the target audio signal after the intensity change processing can be converted into an output audio and output, so as to enable the user to obtain the audio of the target gaze object.
[0254] Further, in order to improve the smoothness of switching, when the target gaze object is switched, a cross-fade strategy can be adopted: gradually reducing the volume of the previous target audio and enhancing the new target voice to avoid sudden changes in hearing.
[0255] Based on any of the above embodiments, after step S140, the voice processing method further comprises step S180.
[0256] Step S180: performing type conversion processing on the target audio signal of the target gaze object.
[0257] Considering specific application scenarios and user needs, after the target audio signal of the target gaze object is identified, type conversion processing can be further performed on the target audio signal of the target gaze object. The type conversion processing includes but is not limited to translation, audio to text, and audio to sign language, etc.
[0258] In this embodiment, through type conversion processing, the audio signal can be converted into different types, and the speaking content of the target gaze object can be displayed in different forms, realizing multi-modal output. This not only improves the user experience, but also widens the application range.
[0259] In a specific embodiment, step S180 can include any one of the following steps S181, step S182 and step S183.
[0260] Step S181: converting the target audio signal of the target gaze object into an audio signal in another language.
[0261] As an implementation, the target audio signal can be converted into an audio signal in another language to translate the speaking content of the target gaze object.
[0262] After obtaining the target audio signal of the target gaze object, it is further detected whether the language of the target audio signal is a preset language, wherein the preset language is a preferred language set by the user in advance, for example, Chinese. The detection method of the language can be: analyzing the acoustic features of the audio of the target gaze object, and determining whether it is a preset speech type according to the analysis result. Or, a deep learning model is used to detect the language of the target audio signal to obtain a detection result, and then it is determined whether it is a preset language according to the detection result. The deep learning model includes but is not limited to: Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), and transformer model.
[0263] When it is detected that the language of the target audio signal is not the preset language, the target audio signal can be translated into an audio signal in the preset language, and then the audio in the preset language is output. Through the conversion of the language, the user can quickly and effectively obtain the speaking content of the target gaze object.
[0264] Step S182, converting the target audio signal of the target gaze object into text.
[0265] As an implementation manner, the target audio signal of the target gaze object can be converted into text, that is, the speaking content of the target gaze object is displayed in the form of text. Specifically, a language recognition model can be used to perform speech recognition on the target audio signal of the target gaze object to obtain a recognition result, and then the recognition result is displayed in the form of text.
[0266] Step S183, converting the target audio signal of the target gaze object into sign language.
[0267] As an implementation manner, the target audio signal of the target gaze object can be converted into sign language, that is, the speaking content of the target gaze object is displayed in the form of sign language. In this case, the deaf-mute person can quickly and effectively obtain the speaking content of the target gaze object.
[0268] Specifically, a language recognition model can be used to perform speech recognition on the target audio signal of the target gaze object to obtain a recognition result, and then the recognition result is converted into sign language and displayed.
[0269] The application scenario of the speech processing method provided by the application will be illustrated below.
[0270] Application scenario 1: AR conference assistant. When a user wears XR glasses to attend a round-table meeting, the camera and microphone on the glasses capture the video and audio of the meeting. Suppose there are multiple participants discussing issues at the same time, the user may have difficulty distinguishing the speech of all participants in the traditional case. The invention first identifies the position and face of each participant through the camera. User A initially focuses on colleague B who is speaking, the system identifies that user A's head is directed at B, so B is set as the current target speaker, and the speech of B is played clearly through the speech separation model. At this time, another colleague C starts to speak, user A turns his head to face C. The system detects that user A's head is directed away from the original direction of B, and the gaze at C lasts more than 300 ms, so it automatically switches the target speaker to C, and the model quickly separates the voice of C and plays it to user A. The whole switching process is smooth and unconscious, without any operation of the user. Even if B and C may occasionally interrupt each other, the fusion model can mainly retain the voice of the current target C, so that user A can hear clearly. During the meeting, if user A is distracted and looks around, due to the setting of the deviation threshold and gaze confirmation, the object will not be easily switched, and the last object of attention will still be locked. In this way, user A can always hear the voice of the main speaker clearly and is not disturbed by noise, greatly improving the efficiency of meeting communication.
[0271] Application scenario 2: hearing aid in noisy environment. A user with mild hearing impairment wears XR glasses supporting the invention and attends a busy family gathering. The room is full of noise and background music. The user starts the auxiliary hearing mode through voice command, and the glasses interface prompts that the target voice enhancement function has been activated. The user sees a friend D not far away and turns his head to D. The system immediately identifies D as the target object, the voice of the speaker is isolated and amplified, and the subtitle display is automatically started: the user's eyes appear the text transcription of D's speech. If D speaks a language that the user is not familiar with, the system will also display / announce the translation of the text into the user's mother tongue. During this period, if the user turns to another person E, the system will also quickly respond to switch the target, enhance the voice of E and display the subtitle. During the whole process, the noisy background laughter and music are significantly lowered, and the user almost only hears the voice of the person he looks at, greatly reducing the auditory burden. Other people in the room do not need to wear any devices, and the system only processes the objects of interest of the user. This application example shows the value of the invention for hearing aid and cross-language communication, and the user can communicate with others in a noisy and multi-lingual environment.
[0272] Application scenario 3: law enforcement and security monitoring. In a security monitoring center, staff wear AR devices to view live video on site and need to listen to the conversations of certain suspicious people to obtain key information. Traditional monitoring requires a directional microphone or manual adjustment, but the application can automatically focus on the target according to the staff's line of sight. For example, there are three suspicious people talking at the same time in the monitoring picture. The staff observes one group of people, and the head points to group A. The system immediately extracts the voices of the people talking in group A and increases the volume, and the voices of the other people are suppressed. When the staff's attention shifts to another group of people, group B, the system seamlessly switches to the voices of the people in group B. This way ensures that the conversation of the target person can be clearly obtained in a complex environment, improving the efficiency of law enforcement and evidence collection. In addition, since the system outputs pure target voice signals, the back end can also record them or convert them into text through voice recognition for archiving as evidence. The application of the application in the security field shows that the head movement driven selection + audio and video separation scheme can enable the operator to hear six ways and see eight directions, and more accurately obtain the required audio information when monitoring multiple tasks.
[0273] As can be seen from the above examples, the method of the application is suitable for various scenarios that require extracting specific speech from noisy, multi-source environments, including but not limited to business meetings, social gatherings, medical consultations, education and training, security monitoring, etc. In these applications, the application can provide unprecedented convenience and functional improvements with its low hardware requirements and high intelligent automation.
[0274] The speech processing device provided by the application is described below. The speech processing device described below can be referred to in conjunction with the speech processing method described above.
[0275] Figure 5 is a structural schematic diagram of the speech processing device provided by the application, as Figure 5 shown, the device includes a first acquisition module 510, an object determination module 520, a second acquisition module 530, and an audio and video processing module 540; wherein:
[0276] The first acquisition module 510 is configured to acquire a current scene image and current head posture data of a user.
[0277] The object determination module 520 is configured to determine a target gaze object according to the current scene image and the current head posture data.
[0278] The second acquisition module 530 is configured to acquire real-time audio data and real-time video data of the target gaze object.
[0279] The audio and video processing module 540 is used to process the real-time audio data and the real-time video data through a speech separation model to determine the target audio signal of the target gaze object.
[0280] The speech processing device provided by this invention acquires the current scene image and the user's current head posture data to select and determine the target gaze object from the scene image seen by the user's head orientation. This method of target object determination relies solely on the XR device's built-in camera and IMU sensor, and compared to manually selecting the target speaker or relying on pre-set parameters, it better aligns with natural human communication habits, enhancing the immersiveness and convenience of the interactive experience. Then, using the XR device's built-in camera and microphone, real-time audio and video data of the target gaze object are acquired. A speech separation model is then used to process the real-time audio and video data to determine the target audio signal of the target gaze object. By fusing visual and audio information, the speech of the target speaker can be accurately separated even in noisy environments or complex scenarios with multiple people speaking simultaneously. In summary, this invention has low hardware requirements, requiring only a single camera and microphone built into the XR device to achieve directional extraction of the target speaker's speech. It eliminates the need for expensive hardware such as microphone arrays, LiDAR, or eye tracking, and does not rely on prior information about the target speaker, making it easy to integrate and deploy on existing XR devices.
[0281] It should be noted that the speech processing device provided in this embodiment of the invention can implement all the method steps implemented in the speech processing method embodiment and achieve the same technical effect. Therefore, the parts and beneficial effects that are the same as those in the method embodiment will not be described in detail here.
[0282] Figure 6 An example is a schematic diagram of the physical structure of an XR device, such as... Figure 6 As shown, the XR device may include a processor 610, a communications interface 620, a memory 630, and a communication bus 640. The processor 610, communications interface 620, and memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions stored in the memory 630 to execute the aforementioned voice processing method.
[0283] In addition, the logic instructions in the memory 630 described above can be implemented in the form of software function units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0284] The device embodiments described above are only schematic, wherein the units illustrated as separate components can or can not be physically separate, and the components illustrated as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0285] From the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and necessary general hardware platform, and of course can also be realized by hardware. Based on such understanding, the above technical solutions essentially or the parts that contribute to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.
[0286] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A voice processing method, characterized by, Applied to an extended reality (XR) device, comprising: obtaining a current scene image and current head pose data of a user; determining a target gaze object according to the current scene image and the current head pose data; obtaining real-time audio data and real-time video data of the target gaze object; processing the real-time audio data and the real-time video data through a voice separation model to determine a target audio signal of the target gaze object; the determining a target gaze object according to the current scene image and the current head pose data comprises: performing face detection on the current scene image through a face detection algorithm to obtain a candidate object and a pixel coordinate thereof; establishing a coordinate system with the user's head as the origin, the front of the user's head as the X-axis, and the vertical upward direction of the origin as the Z-axis, and determining the Y-axis according to the X-axis and the Z-axis; determining the relative coordinates of the candidate object in the coordinate system according to the pixel coordinates and the current scene image; calculating a face direction vector of the candidate object according to the relative coordinates and a reference coordinate of the user's head; calculating a head orientation vector of the user according to the current head pose data; calculating the included angle between the head orientation vector and the face direction vector, and determining the target gaze object from the candidate object according to the calculated included angle.
2. The voice processing method of claim 1, wherein, after the determining a target gaze object according to the current scene image and the current head pose data, further comprising: obtaining a first head pose sequence data of the user, and determining a current reference trajectory according to the first head pose sequence data; obtaining a second head pose sequence data of the user in real time, and detecting whether the gaze object of the user changes according to the second head pose sequence data and the current reference trajectory; if it is detected that the gaze object of the user does not change, then continue to execute: obtaining real-time audio data and real-time video data of the target gaze object; if it is detected that the gaze object of the user changes, then re-execute: obtaining a current scene image and current head pose data of a user.
3. The voice processing method of claim 2, wherein, the obtaining a second head pose sequence data of the user in real time, and detecting whether the gaze object of the user changes according to the second head pose sequence data and the current reference trajectory comprises: obtaining a second head pose sequence data of the user in real time, and determining a real-time motion trajectory according to the second head pose sequence data; encoding the real-time motion trajectory into a real-time feature vector, and encoding the current reference trajectory into a current reference feature vector; calculating a deviation degree according to the real-time feature vector and the current reference feature vector; detecting whether the deviation degree is greater than a preset deviation threshold to detect whether the gaze object of the user changes.
4. The voice processing method of claim 3, wherein, the detecting whether the deviation degree is greater than a preset deviation threshold to detect whether the gaze object of the user changes comprises: detecting whether the deviation degree is greater than a preset deviation threshold; if it is detected that the deviation degree is less than or equal to the preset deviation threshold, then determining that the gaze object of the user does not change; If the deviation is greater than the preset deviation threshold, a timer is started, and it is monitored whether the deviation is greater than the preset deviation threshold within a preset time; If it is monitored that the deviation is greater than the preset deviation threshold within the preset time, it is determined that the gaze object of the user is changed.
5. The voice processing method of claim 1, wherein, The processing of the real-time audio data and the real-time video data by the speech separation model to determine the target audio signal of the target gaze object comprises: The real-time audio data is processed by a time sliding window, and the real-time video data is processed by a frame sliding window; The real-time audio data processed by the time sliding window and the real-time video data processed by the frame sliding window are input into the speech separation model to obtain the target audio signal corresponding to the target gaze object output by the speech separation model; The target audio signal of the target gaze object is obtained by splicing the block audio signals through a smoothing window.
6. The voice processing method of claim 5, wherein, The speech separation model comprises an audio encoder, a visual encoder, an audio-video fusion module, a mask estimation network, and an audio decoder. The real-time audio data processed by the time sliding window and the real-time video data processed by the frame sliding window are input into the speech separation model to obtain the target audio signal corresponding to the target gaze object output by the speech separation model. The audio features output by the audio encoder are obtained by inputting the real-time audio data processed by the time sliding window into the audio encoder. The face visual features output by the visual encoder are obtained by inputting the real-time video data processed by the frame sliding window into the visual encoder. The audio-video fusion features output by the audio-video fusion module are obtained by inputting the audio features and the face visual features into the audio-video fusion module. The time-frequency mask output by the mask estimation network is obtained by inputting the audio-video fusion features into the mask estimation network.
7. The voice processing method of any one of claims 1-6, wherein, The target audio signal corresponding to the target gaze object output by the audio decoder is obtained by inputting the time-frequency mask into the audio decoder. After the processing of the real-time audio data and the real-time video data by the speech separation model to determine the target audio signal of the target gaze object, the following steps are further included: The target audio signal of the target gaze object in the real-time audio data is processed for intensity change; or 8. A speech processing device, characterized by The target audio signal of the target gaze object is processed for type conversion. The first acquisition module is configured to acquire a current scene image and current head posture data of a user. The object determination module is configured to determine a target gaze object according to the current scene image and the current head posture data. The second acquisition module is configured to acquire real-time audio data and real-time video data of the target gaze object. The audio-video processing module is configured to process the real-time audio data and the real-time video data by a speech separation model to determine a target audio signal of the target gaze object. The object determination module is specifically configured to: A face detection algorithm is used to perform face detection on the current scene image to obtain candidate objects and their pixel coordinates. A coordinate system is established with a head of a user as an origin, a front of the head of the user as an X axis, and a vertical upward direction of the origin as a Z axis, and a Y axis is determined according to the X axis and the Z axis; According to the pixel coordinates and the current scene image, relative coordinates of the candidate object in the coordinate system are determined; According to the relative coordinates and reference coordinates of the head of the user, a face direction vector of the candidate object is calculated; According to the current head posture data, a head orientation vector of the user is calculated; An included angle between the head orientation vector and the face direction vector is calculated, and a target gaze object is determined from the candidate object according to the calculated included angle.
9. An XR device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, wherein, The processor implements the voice processing method in any one of claims 1 to 7 when executing the computer program.
Citation Information
Patent Citations
Interaction method and device in virtual reality scene, equipment and storage medium
CN116048281A
Target speaker voice extraction method and system based on sight tracking technology
CN117877494A