Separate sound projection method and system
Patent Information
- Application Number
- CN202611086176.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-22
AI Technical Summary
但如果仅仅依据固定的左右声道到达时间差阈值来判断声像位置是否合理,就会脱离音箱摆放距离这一关键变量,当用户改变摆放后,原有调节策略与实际听音几何关系不再对应,致使调节后的声源方位依然无法与画面口型重合
本发明公开了一种分离式音响投影方法及系统,解决了传统立体声播放中声音方位与画面人物口型位置脱离的问题。本发明通过获取左右音箱相对于画面中心线的横向摆放位置,计算左右声道信号到达预设听音位置的第一时间差,并据此识别当前声像在立体听音空间中的初始落点。同时对同步播放的视频画面进行口型识别,提取人物面部成像坐标及口型位置的横向偏移与纵向高度,结合预设的视听映射关系确定画面口型应当发声的目标声像落点。当初始声像落点与目标声像落点在立体听音空间中出现分离时,本发明通过立体声声像定位技术计算使声像方位从初始位置迁移至目标位置所需的声道间时间差调整量,并对先到达的声道信号施加相应延时处理,使处理后的左右声道信号到达预设听音位置的时间差精确匹配目标声像方位。该方法实现了声音方位与画面口型的动态贴合,提升了视听一体化的沉浸体验效果。
Smart Images

Figure CN122802845A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information technology, and in particular to a separated sound projection method and system. BACKGROUND
[0002] Home audio-visual entertainment is becoming an important part of modern life, and projection systems are widely welcomed for their immersive experience with large screens. In a separated sound projection system, in order to create a wider sound field effect, users often need to move the left and right sound boxes to the sides, which can significantly enhance the spatial surround effect. However, this approach can cause the sound localization to deviate from the picture content, especially when the characters in the picture are speaking, the sound direction heard by the audience comes from the sound box position, rather than the picture central area where the character's mouth is located, forming a sound image separation phenomenon, which seriously damages the realism of the viewing experience. Existing sound image adjustment schemes mostly rely on fixed sound channel delay parameters or preset sound field modes. These methods may be effective when the sound box position is fixed, but in the case of flexible adjustment of the sound box placement by the user according to the room layout, since the pre-set delay parameter still takes effect according to the original sound box distance, the delay amount of the sound channel signal no longer matches the actual placement, and the sound source direction perceived by the audience deviates from the picture center, and the sound image separation phenomenon reappears and cannot be eliminated. When the horizontal distance between the sound boxes changes, the time difference between the left and right sound channel signals reaching the human ear will also change, and the human ear determines the sound source direction based on this time difference. Once the placement distance changes, the time difference reference for determining the sound source direction is no longer valid. However, if the sound image position is determined only based on the fixed time difference threshold between the left and right sound channels, it will deviate from the key variable of sound box placement distance. When the user changes the placement, the original adjustment strategy no longer corresponds to the actual listening geometry, causing the adjusted sound source direction to still not coincide with the picture mouth shape. For example, when the sound boxes are moved one meter outward, even if the left and right sound channels are played synchronously, the sound source perceived by the audience will still be scattered on both sides, rather than concentrated on the picture central character's mouth. At this time, if the original time difference reference is still used, it will miss the sound image that has already been separated, and will not correct the playback segment that needs to be corrected. As can be seen, the sound image positioning reference deviates from the actual sound box placement state, causing the sound source direction determined by the reference to remain at the original sound box distance, while the actual sound source direction perceived by the audience deviates to both sides as the sound boxes move outward. A fixed pointing deviation is formed between them, i.e. the reference always determines the sound source to be in the picture center, while the actual sound source has deviated from the center, and this deviation increases as the sound boxes move outward and cannot be self-corrected. Therefore, how to dynamically establish a sound image positioning reference according to the actual placement of the sound boxes, and accurately identify the time period when the sound direction deviates from the visual focus in combination with the real position information of the characters' mouths in the picture, has become a key problem for achieving sound and picture synchronization and improving immersive viewing experience. SUMMARY
[0003] This invention provides a method for split-type sound projection, mainly including: Obtain the lateral placement of the first speaker located on one side of the center line of the screen and the second speaker located on the other side of the center line of the screen relative to the center line of the screen, and calculate the first time difference between the left channel signal and the right channel signal reaching the preset listening position. Based on the first time difference, the sound source location of the left and right channel signals is identified to determine the initial sound image landing point in the stereo listening space after the left and right channel signals are synthesized. Lip-reading recognition is performed on the faces of people in video images that are played synchronously with the left and right channel signals to identify the position of the lip movements of the person currently speaking. The planar orientation of the lip movements in the image is obtained on the projected image. Based on the planar orientation and the preset audiovisual mapping relationship, the target sound image landing point where the lip movements in the image should be pronounced is determined. Evaluate whether the initial sound image landing point and the target sound image landing point are separated from each other in the stereo listening space, and obtain the sound image separation time period; For the audio signal within the time period of the separation of the sound image, the time difference adjustment amount between the channels is determined based on the horizontal placement position, the initial sound image orientation and the target sound image orientation, and is used as the second time difference; Based on the second time difference, a delay is applied to the channel signal that arrives at the preset listening position first in the audio signal to obtain the output audio signal and send it to the first speaker and the second speaker.
[0004] Furthermore, the step of obtaining the lateral placement positions of the first speaker located on one side of the center line of the image and the second speaker located on the other side of the center line of the image relative to the center line of the image, and calculating the first time difference between the left channel signal and the right channel signal reaching the preset listening position, includes: deploying infrared ranging sensors from the vertical line of the center of the projected image to both sides to obtain the lateral placement positions of the first speaker and the second speaker relative to the vertical line of the center; forming a right-angled triangle in the horizontal plane with the lateral placement positions and the longitudinal front-back distance from the preset listening position to the vertical line of the center, and using the Pythagorean theorem to calculate the path lengths of the first speaker and the second speaker to the preset listening position; dividing the difference between the two path lengths from the first speaker and the second speaker to the preset listening position and the second speaker to the preset listening position by the speed of sound to obtain the first time difference.
[0005] Furthermore, the step of identifying the sound source location of the left and right channel signals based on the first time difference and determining the initial sound image landing point in the stereo listening space after the left and right channel signals are synthesized includes: determining the offset direction of the synthesized sound image relative to the center line of the preset listening position based on the sign of the first time difference; establishing a horizontal coordinate system of the stereo listening space with the preset listening position as the origin and the center line as the vertical axis; calculating the azimuth angle of the synthesized sound image relative to the vertical axis using an arcsine function based on the ratio of the absolute value of the first time difference to the interaural distance; marking the azimuth point along the vertical axis with a preset sensing radius as the distance and the azimuth angle as the subtended angle in the offset direction to obtain the initial sound image landing point.
[0006] Furthermore, the step of performing lip-syncing on the face of a person in a video frame played synchronously with the left and right channel signals to identify the lip position of the person currently speaking includes: extracting a video frame sequence frame by frame from the video frame played synchronously with the left and right channel signals; performing face detection on each frame using a multi-task cascaded convolutional neural network to obtain a face detection box; for the face image of the person within the face detection box, using a face key point localization method based on a depth alignment network to mark the coordinates of key points on the upper lip contour, lower lip contour, and corners of the mouth; determining the opening and closing amplitude of the lip region by the vertical pixel distance between the central key point of the upper lip and the central key point of the lower lip; when the difference in opening and closing amplitude between adjacent video frames is greater than a preset speaking threshold, it is determined as a speaking frame and retained; otherwise, it is determined as a silent frame and discarded; for the retained speaking frames, the centroid of the region enclosed by the central key points of the upper lip, central key points of the lower lip, and corners of the mouth is used as the lip coordinates, and aligned synchronously with the left and right channel signals according to the video frame timestamp to obtain the lip position of the image.
[0007] Furthermore, the step of obtaining the planar orientation of the lip-sync position on the projected screen, and determining the target sound image landing point where the lip-sync should be emitted based on the planar orientation and a preset audiovisual mapping relationship, includes: obtaining a lateral offset ratio with positive and negative signs from the offset of the lateral pixel component relative to the center vertical line, and obtaining a vertical height ratio from the height of the vertical pixel component relative to the bottom edge, as the planar orientation; using the preset audiovisual mapping relationship, converting the lateral offset ratio into the target azimuth angle, and converting the vertical height ratio into the target elevation angle; and marking the target sound image landing point along a preset perception radius direction according to the target azimuth angle and the target elevation angle, with the preset listening position as the origin and the direction directly in front of the center line as the reference.
[0008] Furthermore, the method also includes: evaluating the degree to which the lip-sync point on the projected image is positioned to the left or right within the visible area of the projected image, thereby obtaining the planar orientation of the lip-sync point on the projected image.
[0009] Furthermore, the step of evaluating whether the initial sound image landing point and the target sound image landing point are separated from each other in the stereo listening space to obtain the sound image separation time period includes: calculating the Euclidean distance between the initial sound image landing point and the target sound image landing point in the stereo listening space based on the coordinates of the initial sound image landing point and the target sound image landing point; marking the current sound frame as a separation frame when the Euclidean distance is greater than a preset separation threshold, and marking it as a fit frame otherwise; according to the timestamps corresponding to the separation frames, grouping multiple separation frames with an interval of less than twice the video frame period into the same time period, taking the earliest separation frame timestamp as the start point of the time period and the latest separation frame timestamp as the end point of the time period; and not including the fit frames in the time period merging to obtain the sound image separation time period.
[0010] Furthermore, for the audio signal within the time period of the sound image separation, based on the lateral placement position, the initial sound image orientation, and the target sound image orientation, the adjustment amount of the inter-channel time difference is located as the second time difference, including: extracting the target azimuth angle of the synthesized sound image relative to the centerline from the target sound image landing point; using the sine correspondence between the binaural time difference and the azimuth angle, and back-calculating the target binaural time difference corresponding to the target azimuth angle based on the binaural distance and sound velocity values; determining the initial azimuth angle from the initial sound image landing point, and back-calculating the initial binaural time difference using the same sine correspondence, wherein the initial binaural time difference is consistent with the first time difference; subtracting the initial binaural time difference from the target binaural time difference, and combining it with the signal arrival order caused by the lateral placement position, to obtain the second time difference.
[0011] Furthermore, the step of applying a delay to the channel signal that arrives at the preset listening position first in the audio signal according to the second time difference to obtain an output audio signal and sending it to the first speaker and the second speaker includes: determining the channel that arrives at the preset listening position first from the left channel signal and the right channel signal according to the sign of the second time difference; multiplying the absolute value of the second time difference by the sampling rate of the audio signal and rounding down to obtain the number of delay samples; applying a sample delay consistent with the number of samples to the channel that arrives first using a digital delay buffer; keeping the other channel signal in its original timing; and after the delay, sending the left channel signal to the first speaker amplifier channel and the right channel signal to the second speaker amplifier channel to obtain the output audio signal.
[0012] This application provides a split-type audio projection system, the system comprising: The first acquisition module is used to acquire the horizontal placement positions of the first speaker located on one side of the center line of the screen and the second speaker located on the other side of the center line of the screen relative to the center line of the screen, and to calculate the first time difference between the left channel signal and the right channel signal reaching the preset listening position. The first calculation module is used to identify the sound source location of the left and right channel signals based on the first time difference, and determine the initial sound image landing point of the synthesized left and right channel signals in the stereo listening space. The lip-reading recognition module is used to recognize the lip movements of people in a video frame that is played synchronously with the left and right channel signals, and to identify the position of the lip movements of the person who is currently speaking. The target sound image determination module is used to obtain the planar orientation of the lip shape position on the projected screen, and determine the target sound image landing point where the lip shape should be uttered based on the planar orientation and the preset audiovisual mapping relationship. The evaluation module is used to evaluate whether the initial sound image landing point and the target sound image landing point are separated from each other in the stereo listening space, and to obtain the sound image separation time period. The positioning module is used to locate the time difference adjustment between the sound channels for the audio signal within the time period of the sound image separation, based on the horizontal placement position, the initial sound image orientation and the target sound image orientation, as the second time difference; The delay processing module is used to apply delay processing to the channel signal that arrives at the preset listening position first in the audio signal according to the second time difference, so as to obtain the output audio signal and send it to the first speaker and the second speaker.
[0013] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects: This invention discloses a method and system for split-sound projection, solving the problem of sound location deviating from the lip-sync position of characters in traditional stereo playback. The invention obtains the lateral placement of the left and right speakers relative to the center line of the screen, calculates the first time difference between the left and right channel signals arriving at the preset listening position, and identifies the initial landing point of the current sound image in the stereo listening space. Simultaneously, lip-sync recognition is performed on the synchronously played video image, extracting the facial imaging coordinates and the lateral offset and vertical height of the lip-sync position. Combined with a preset audiovisual mapping relationship, the target sound image landing point where the lip-sync should produce sound is determined. When the initial sound image landing point and the target sound image landing point separate in the stereo listening space, this invention uses stereo sound image localization technology to calculate the inter-channel time difference adjustment required for the sound image location to migrate from the initial position to the target position, and applies a corresponding delay to the first arriving channel signal, ensuring that the time difference between the processed left and right channel signals arriving at the preset listening position precisely matches the target sound image location. This method achieves dynamic alignment between sound location and lip-sync, enhancing the immersive audiovisual experience. Attached Figure Description
[0014] Figure 1 This is a flowchart of the split-type sound projection method of the present invention.
[0015] Figure 2This is a schematic diagram of the split-type sound projection method of the present invention.
[0016] Figure 3 This is another schematic diagram of the split-type sound projection method of the present invention.
[0017] Figure 4 This is a schematic diagram of the structure of the split-type audio projection system of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and thoroughly described below with reference to the accompanying drawings. The described embodiments are merely some embodiments of the present invention.
[0019] like Figures 1-4 The separate audio projection method in this embodiment may specifically include: Step S101: Obtain the horizontal placement positions of the first speaker located on one side of the center line of the screen and the second speaker located on the other side of the center line of the screen relative to the center line of the screen, and calculate the first time difference between the left channel signal and the right channel signal reaching the preset listening position.
[0020] Infrared ranging sensors are deployed to the left of the center vertical line of the projected image to obtain the lateral placement position of the first speaker from the center vertical line, and infrared ranging sensors are deployed to the right to obtain the lateral placement position of the second speaker from the center vertical line. Based on the lateral placement positions of the first and second speakers and the longitudinal distance from the preset listening position to the center vertical line, a right triangle is formed in the horizontal plane, with the lateral placement position as one leg and the longitudinal distance as the other leg. The Pythagorean theorem is used to calculate the first path length from the first speaker to the preset listening position and the second path length from the second speaker to the preset listening position. The difference between the first and second path lengths is divided by an empirically derived value of the speed of sound in dry air at room temperature to obtain the first time difference between the left and right channel signals reaching the preset listening position.
[0021] In home audio-visual entertainment scenarios, after the user flexibly adjusts the placement of the first and second speakers of a split audio projection system according to the room layout, the time difference between the channel signals reaching the preset listening position changes with the geometric relationship of the placement, and it is necessary to re-establish the time difference benchmark based on the actual placement.
[0022] In one possible implementation, the central vertical line of the projected image is obtained by extending downwards from the horizontal midpoint of the projected image, serving as the geometric reference line for dividing the left and right audio channels. The first speaker is located to the left of the central vertical line, and the second speaker is located to the right of the central vertical line, with each speaker responsible for the electroacoustic conversion between the left and right channel signals.
[0023] For example, an infrared ranging sensor transmitting and receiving element is fixedly installed below the projection screen along the central vertical line. An infrared beam is emitted to the left to illuminate the side wall of the first speaker's enclosure. The lateral placement position of the first speaker relative to the central vertical line is obtained by measuring the round-trip time of the infrared beam. The side wall of the second speaker's enclosure is illuminated to the right in the same manner to obtain the lateral placement position of the second speaker relative to the central vertical line. The range of the infrared ranging sensor covers the lateral range that speakers in a typical living room may be placed in.
[0024] It should be noted that the preset listening position refers to the center point of the sofa seat where the user usually watches movies. The same infrared ranging sensor's transmitting and receiving elements are arranged along the direction directly opposite to the center vertical line, emitting an infrared beam towards the center point of the sofa seat. The longitudinal distance from the preset listening position to the center vertical line is obtained by measuring the round-trip time of the infrared beam.
[0025] Specifically, the center height of the speaker's tweeter is on the same horizontal plane as the ear height at the preset listening position, and the straight path between the speaker and the listening position lies within this horizontal plane. Using the lateral placement of the first speaker as one leg of a right triangle, and the longitudinal distance from the preset listening position to the central vertical line as the other leg, the right angle between the two legs lies on the central vertical line. The hypotenuse is the first path length from the first speaker to the preset listening position. Using the Pythagorean theorem, the hypotenuse length is the square root of the sum of the squares of the two legs. Substituting the same geometric relationship into the lateral placement of the second speaker and the same longitudinal distance, we obtain the second path length from the second speaker to the preset listening position. When the lateral placement of the first and second speakers changes, the first and second path lengths are updated accordingly.
[0026] In one embodiment, the empirical value for the speed of sound is taken from a commonly used reference value for the propagation speed of sound waves under normal temperature and dry air conditions, which is approximately 340 meters per second. The difference between the first path length and the second path length is divided by the empirical value for the speed of sound to obtain the first time difference between the left channel signal and the right channel signal arriving at the preset listening position. This first time difference is positive or negative; a positive value indicates that the right channel signal arrives at the preset listening position first, and a negative value indicates that the left channel signal arrives at the preset listening position first.
[0027] It is understandable that when the user moves the first speaker and the second speaker outward symmetrically, the path lengths from the two speakers to the preset listening position increase synchronously and remain equal, and the first time difference approaches zero; when the user places the two speakers asymmetrically, the first time difference becomes a non-zero value, and the absolute value increases with the degree of asymmetry.
[0028] For example, in a living room environment, the first speaker is positioned horizontally at a distance of 0.8 meters, the second speaker at a distance of 1.2 meters, and the longitudinal distance from the preset listening position to the center vertical line is 3 meters. The path length is calculated using the Pythagorean theorem, where L represents the path length, x represents the horizontal position of the speaker, and y represents the longitudinal distance. The length of the first path is approximately 3.105 meters, and the length of the second path is approximately 3.231 meters. The difference between the two is approximately 0.126 meters. Dividing this by the speed of sound, 340 meters per second, yields the first time difference of approximately 0.37 milliseconds, which serves as a time difference benchmark reflecting the current placement of the speaker.
[0029] Step S102: Based on the first time difference, identify the sound source location of the left and right channel signals to determine the initial sound image landing point of the current synthesized left and right channel signals in the stereo listening space.
[0030] Based on the sign of the first time difference, the offset direction of the synthesized sound image relative to the center line of the preset listening position is determined. If the first time difference is positive, the synthesized sound image is offset to the left of the center line; otherwise, it is offset to the right. A stereo listening space horizontal coordinate system is established with the preset listening position as the origin, the center line as the vertical axis, and the horizontal direction perpendicular to the center line as the horizontal axis. The azimuth angle of the synthesized sound image relative to the vertical axis is calculated using an arcsine function based on the ratio of the absolute value of the first time difference to the interaural distance. Based on the offset direction and the azimuth angle, azimuth points are marked along the vertical axis in the stereo listening space horizontal coordinate system, with a preset sensing radius as the distance and the azimuth angle as the subtended angle, towards the corresponding offset direction. This yields the initial sound image landing point perceived after the current synthesis of the left and right channel signals.
[0031] The first time difference reflects the timing difference between the left and right channel signals arriving at the preset listening position under the current speaker placement geometry. Next, the sound image perception location corresponding to this timing difference is identified.
[0032] In one possible implementation, the center line of the preset listening position is defined as a vertical line passing through the preset listening position and parallel to the central vertical line, serving as a listening reference line for distinguishing left and right perceptual orientation. When the listener is located at the preset listening position, the left ear faces to the left of the center line, and the right ear faces to the right of the center line.
[0033] It should be noted that the mechanism by which the human ear determines the location of a sound source based on the time difference between the two ears is as follows: when the sound source is biased towards the listener, the ear closer to the sound source receives the sound wave first, while the other ear receives it slightly later. The difference in the arrival times of the two ears is decoded into a horizontal azimuth angle in the auditory cortex. The first time difference is obtained by subtracting the arrival time of the second speaker's signal from the arrival time of the first speaker's signal, with the sign corresponding to the biased side. The first speaker corresponds to the left channel, and the second speaker corresponds to the right channel. The preset listening position is the standard listening point 3 meters behind the midpoint of the line connecting the two speakers. The midline is the intersection of the midsagittal plane directly in front of the listener and the horizontal plane. When the first time difference is positive, it indicates that the left channel signal arrives later than the right channel signal, meaning the right channel signal arrives at the preset listening position first, and the synthesized sound image is biased to the left of the center line. When the first time difference is negative, it indicates that the left channel signal arrives first, and the synthesized sound image is biased to the right of the center line. When the first time difference is zero, the synthesized sound image falls on the center line. The sign of the first time difference indicates the direction of the offset, and its absolute value determines the magnitude of the offset. A stereoscopic listening space horizontal coordinate system is established with the preset listening position as the origin O, the direction directly in front of the center line as the positive vertical axis Y, and the direction perpendicular to the vertical axis and pointing to the listener's right side in the horizontal plane as the positive horizontal axis X. The stereoscopic listening space horizontal coordinate system XOY is located in the horizontal plane where the listener's ears are located, and the sound image landing point is marked with its spatial position in this plane using horizontal and vertical coordinate values.
[0034] In one embodiment, the azimuth conversion process is based on the ratio of the interaural time difference to the interaural distance. First, the geometric relationship between the left and right speakers and the preset listening positions is determined based on the first time difference. Then, the time difference between the two ears receiving the sound waves under this geometric relationship is calculated. Let the absolute value of the interaural time difference be Δt (in seconds); the interaural distance be d (0.175 meters); and the speed of sound be c (340 meters per second). The product of the interaural time difference Δt and the speed of sound c is the interaural path difference, and its ratio to the interaural distance d is denoted as k, i.e., k = Δtc / d. When the absolute value of k is not greater than 1, the ratio k is substituted into the arcsine function to obtain the azimuth angle θ of the synthesized sound image relative to the vertical axis, i.e., θ = arcsink. When the absolute value of k is greater than 1, the sound image azimuth angle is 90 degrees. For example, when the absolute value of the time difference between the two ears is 0.12 milliseconds, Δt is 0.00012 seconds, and Δtc is calculated to be 0.0408 meters. k = 0.0408 / 0.175 ≈ 0.233, and the azimuth angle θ is approximately 13.5 degrees after arcsine conversion.
[0035] Preferably, to mark the specific position of the sound image azimuth point in the horizontal coordinate system, a perception radius R is preset to characterize the estimated scale of the listener's subjective perception of the sound source distance. The value is taken as a reference to the empirical value of the human ear's perception of the distance to a sound source in front in a living room environment, with 1.5 meters as the standard perception distance. The perception radius R is independent of the longitudinal distance from the aforementioned preset listening position to the central vertical line in both value and meaning. The stereo listening space is a three-dimensional sound field space established with the center of the listener's head as the origin, including a horizontal coordinate system and a vertical coordinate system, used to describe the spatial position of the sound source and auditory perception. The central vertical line is the projection of the straight line connecting the midpoints of the left and right speakers to the preset listening position onto the horizontal plane, serving as a reference for the vertical axis of the coordinate system. In the horizontal coordinate system of the stereo listening space, with the vertical axis as the reference, the azimuth angle θ is rotated in the offset direction, i.e., the left and right rotation directions indicated by positive and negative signs, to obtain an azimuth ray emanating from the coordinate origin. Starting from the origin of the coordinate system, advance the distance corresponding to the perception radius R along the azimuth ray direction to obtain an azimuth point located within the horizontal plane of the stereo listening space. This azimuth point is the initial sound image landing point, representing the virtual sound source position perceived by the listener in the horizontal plane after the left and right channel signals are mixed according to the current amplitude ratio. The initial sound image landing point is used in subsequent steps to extract azimuth parameters for matching and correction with the sound source position identified in the video image.
[0036] Understandably, the initial sound image landing point refers to the virtual sound source location perceived by the listener on the horizontal plane after auditory localization based on the time difference and intensity difference of the sound received by both ears. The midline is defined as the direction of the midsagittal plane projection when the listener faces the center of the display screen. When the first and second speakers are symmetrically placed, the first time difference approaches 0, and the azimuth angle θ calculated by the binaural time difference localization algorithm approaches 0 degrees. The initial sound image landing point falls on the vertical axis and is close to the centerline. When the two speakers are asymmetrically placed or moved outwards, causing the sound path on one side to be significantly longer than that on the other side, the time difference of the sound received by both ears increases, and the azimuth angle... As it increases, where c is the speed of sound (340 m / s), Δt is the time difference between the two ears, and d is the distance between the two ears (approximately 0.17 m). The further the initial sound image deviates from the center line, the more it reflects the existing directional deviation between the perceived sound source and the center of the image under the current playback state.
[0037] Step S103: Perform lip-reading recognition on the faces of people in the video frame that is played synchronously with the left and right channel signals, and identify the position of the lip-reading of the person who is currently speaking.
[0038] Video frame sequences are extracted frame by frame from the video footage played synchronously with the left and right channel signals. A multi-task cascaded convolutional neural network is used to perform face detection on each frame, obtaining a bounding box containing the person's face in the image. For the face image within the bounding box, a facial landmark localization method based on a depth alignment network is used to mark the coordinates of the upper lip contour, lower lip contour, and the corners of the mouth. The opening and closing amplitude of the lip region is determined by the vertical pixel distance between the central landmark of the upper lip and the central landmark of the lower lip. Based on the changes in the opening and closing amplitude of the lip region in the continuous video frame sequence, if the difference in opening and closing amplitude between adjacent video frames, expressed in pixels, is greater than a preset vocalization threshold, it is determined to be a vocal frame and retained; otherwise, it is determined to be a silent frame and discarded. For the retained vocal frames, the centroid of the region enclosed by the central landmark of the upper lip, the central landmark of the lower lip, and the corners of the mouth is used as the lip shape coordinates. These coordinates are then aligned synchronously with the left and right channel signals according to the video frame timestamp to obtain the lip shape position of the person currently speaking.
[0039] In home audio-visual entertainment scenarios, the real-time spatial position of the mouth shape of a person speaking in the projected image constitutes the visual anchor point for subsequent sound-image matching. The mouth shape position in the image is extracted and located from the video image played synchronously with the left and right channel signals.
[0040] In one possible implementation, the video frame rate is typically 24 to 60 frames per second. Video frames are captured frame-by-frame from the decoding output of the projection playback channel, resulting in a chronologically ordered sequence of video frames. The projection playback channel refers to the software data channel from the video decoder to the projection device, responsible for transmitting the decoded video data to the projection display module. Each video frame corresponds to a static image with a fixed pixel resolution, and frames are linked by a frame number and a timestamp. The playback clock uses the system audio clock as the reference clock source. The timestamps of the video frames and the sampling timestamps of the left and right channel signals are both generated based on this clock. Audio and video are synchronized by comparing and calibrating the video frame timestamps with the audio sampling timestamps. When the difference between the two timestamps exceeds 33.3 milliseconds, a buffer adjustment is triggered to eliminate latency.
[0041] It should be noted that the multi-task cascaded convolutional neural network is a classic face detection method, consisting of three convolutional networks connected in series. Each preceding network coarsely filters candidate regions, while subsequent networks refine the position and size of these regions, outputting bounding boxes for all faces in the image and five coarse localization keypoints. Each frame in the video frame sequence is used as input to this network, outputting the face detection boxes containing the faces of the individuals. When multiple individuals are present in the frame, multiple face detection boxes are output, each recorded with its top-left corner pixel coordinates and width and height pixel dimensions.
[0042] Preferably, considering that characters in film and television scenes are at different distances from the foreground, the face detection boxes are sorted from largest to smallest by pixel area. The face detection boxes at the top correspond to the main characters in the scene, and subsequent key point localization and voice determination prioritize the face detection boxes at the top.
[0043] Specifically, the depth alignment network uses a cascaded convolutional neural network for facial landmark regression. This embodiment uses a structure containing three levels of networks, with each level progressively refining the landmark positions. The facial image captured within the face detection box is scaled to 112×112 pixels and then fed into the network, outputting the pixel coordinates of 68 landmarks on the face. Of these 68 landmarks, the upper lip contour is described by 7 landmarks, with the one located in the exact center being the upper lip center landmark; the lower lip contour is described by 7 landmarks, with the one located in the exact center being the lower lip center landmark; additionally, landmarks at the left and right corners of the mouth are included. The upper lip center landmark is located in the exact center of the upper lip contour, and the lower lip center landmark is located in the exact center of the lower lip contour.
[0044] In one embodiment, the opening and closing amplitude of the lip region is defined as the vertical pixel distance between the central key point of the upper lip and the central key point of the lower lip, where vertical refers to the vertical direction along the video frame. Let the vertical pixel coordinate of the central key point of the upper lip be y1, and the vertical pixel coordinate of the central key point of the lower lip be y2, then the opening and closing amplitude h = y2 - y1, in pixels. When the person's mouth is closed, h takes a smaller value, and when the person's mouth is open and making a sound, h increases significantly. Further, the opening and closing amplitudes calculated frame by frame in a continuous video frame sequence are arranged according to the frame number to form a curve of the opening and closing amplitude changing over time. Let the opening and closing amplitude of the current frame be hcurrent, and the opening and closing amplitude of the previous frame be hprevious, and the change in opening and closing amplitude between adjacent video frames be denoted as d. The calculation rule is to first take the absolute value of the value obtained by subtracting hprevious from hcurrent and then assign it to d, in pixels. The change in opening and closing amplitude d represents the instantaneous intensity of the lip opening and closing activity between two adjacent frames, without distinguishing between opening and closing actions. The preset vocal threshold is in pixels and is denoted as h0. The value is taken from the minimum recognizable change in the opening and closing of a person's lips when speaking.
[0045] For example, the value is 3 pixels in a high-definition image with a resolution of 1920 x 1080, and 6 pixels in an ultra-high-definition image with a resolution of 3840 x 2160.
[0046] Specifically, if the change in the opening and closing amplitude d is greater than the preset sound threshold h0, the frame is determined to be a sound frame and retained; if d is less than or equal to h0, the frame is determined to be a silent frame and removed from subsequent processing. The sound frame corresponds to the frame in the picture where a person is opening and closing their lips to produce speech, and the silent frame corresponds to the frame where a person's lips are closed or have almost no opening and closing motion. This determination distinguishes between the sound-producing and non-sound-producing periods in the picture, avoiding the misidentification of the lip shape position during silent periods as the sound image matching target.
[0047] In one embodiment, the acquired audio data is in stereo two-channel format, including left channel and right channel signals. For the retained sound frames, the arithmetic mean of the pixel coordinates of four key points—the center key point of the upper lip, the center key point of the lower lip, the left corner key point, and the right corner key point—is taken. The horizontal coordinate is the sum of the four horizontal coordinates divided by 4, and the vertical coordinate is the sum of the four vertical coordinates divided by 4. The resulting coordinate point is the centroid of the area enclosed by the four key points, and is used as the lip shape coordinate. The lip shape coordinate is a two-dimensional coordinate point on the pixel plane, refreshed frame by frame as the sound frames progress. Further, the timestamp of each sound frame is used as an alignment anchor point, aligned with the sampling clock of the left and right channel signals. The left and right channel signals are arranged according to sampling points, each sampling point having a sampling time. Audio sampling points whose sampling time falls between the timestamp of one sound frame and the timestamp of the next sound frame are assigned to the lip shape coordinates corresponding to that sound frame. Since video frame rates are typically 25 to 30 frames per second, while audio sampling rates are 44,100 Hz, and single-frame durations are approximately 33 to 40 milliseconds, the changes in human lip movements are relatively gradual within this duration. By uniformly assigning inter-frame audio sampling points to the lip movement coordinates of the previous frame, audiovisual consistency can be maintained when driving virtual avatar lip-sync, meeting the accuracy requirements of real-time interactive scenarios. This forms a sequence of lip movement positions on the screen, strung together by timestamps. Each element in the sequence represents the lip movement coordinates on a pixel plane and is bound to an audio time interval that is synchronized with it.
[0048] Understandably, when only one person speaks in the scene at any given moment, the aforementioned sound frame determination directly outputs the lip position of that person. When multiple people appear in the scene at the same time but only one person speaks, the change in the opening and closing amplitude d is calculated for each of the multiple face detection boxes. Only the face detection box corresponding to the person whose change in the opening and closing amplitude d exceeds the preset sound threshold h0 is judged as the face box of the speaking subject, and the face detection boxes corresponding to the other people are judged as the face boxes of the silent subject. The lip position is taken from the centroid of the lip area within the face box of the speaking subject to obtain the lip position of the current person speaking.
[0049] Step S104: Obtain the planar orientation of the lip-sync position on the projected screen. Based on the planar orientation and the preset audiovisual mapping relationship, determine the target sound image landing point where the lip-sync should be emitted in the stereoscopic listening space where the preset listening position is located.
[0050] The horizontal and vertical pixel components are extracted from the pixel coordinates of the lip-sync position on the screen. Using the center vertical line of the projected image as the horizontal zero point and the bottom edge of the projected image as the vertical zero point, the horizontal pixel component is subtracted from the center vertical line pixel coordinates and then divided by half the total number of pixels in the width of the projected image to obtain a signed horizontal offset ratio. The vertical pixel component is divided by the total number of pixels in the height of the projected image to obtain the vertical height ratio, which serves as the planar orientation of the lip-sync position on the projected image. Using a preset audiovisual mapping relationship, the horizontal offset ratio is multiplied by a preset value for the maximum recognizable azimuth angle on one side in the listener's horizontal plane to obtain a signed target azimuth angle; the vertical height ratio is subtracted by 0.5 and then multiplied by a preset value for the maximum recognizable elevation angle on one side in the listener's vertical plane to obtain the target elevation angle. Using the preset listening position as the origin of the coordinate system and the direction directly in front of the center line as the reference direction, the system rotates in the horizontal plane according to the target azimuth angle and raises or lowers in the vertical plane according to the target elevation angle. The azimuth point is marked along the direction of the perception radius R to obtain the target sound image landing point where the mouth shape in the picture should produce sound.
[0051] In home audio-visual entertainment scenarios, the lip-sync positions on the screen exist as two-dimensional coordinate points on a pixel plane. Next, these lip-sync positions are mapped from the projected screen coordinate system to the stereoscopic listening space coordinate system where the preset listening position is located, thus marking the target sound image landing point where the lip-sync should be emitted. The preset listening position is the center point of the listener's head, serving as the reference origin for sound location perception. The resulting target sound image landing point indicates the spatial location in the stereoscopic listening space where the sound corresponding to the speaker's lip-sync should be presented, allowing the audience to perceive that the sound originates from the speaker's location on the screen, achieving spatial consistency between audio and video synchronization.
[0052] In one possible implementation, the resolution of the projected image is denoted as the total number of pixels in width W and the total number of pixels in height H. The vertical line at the center of the projected image is located at the horizontal pixel coordinate W divided by 2. The horizontal pixel coordinate of the lip shape position is denoted as u, and the vertical pixel coordinate is denoted as v, where v is measured upwards with the bottom edge of the projected image as the zero point.
[0053] Specifically, the rule for determining the lateral offset ratio is as follows: px is the horizontal pixel coordinate of the lip-sync position on the screen. px falls within a closed interval of -1 to 1. A negative px value indicates the lip-sync position is to the left of the center vertical line, and a positive px value indicates the lip-sync position is to the right of the center vertical line. When px is 1, the lip-sync position is flush with the right edge of the projected image; when px is -1, the lip-sync position is flush with the left edge of the projected image; and when px is 0, the lip-sync position is exactly on the center vertical line. The absolute value of px is equal to the ratio of the horizontal distance from the lip-sync position to the center vertical line to half the width of the image. The vertical height ratio is determined by the rule py = v / H, where v is the vertical pixel coordinate of the lip-sync position. py falls within a closed interval of 0 to 1. py close to 0 indicates the lip-sync position is close to the bottom edge of the projected image, and py close to 1 indicates the lip-sync position is close to the top edge of the projected image. px and py together constitute the planar orientation of the lip-sync position on the projected image.
[0054] It should be noted that the audiovisual mapping relationship is a conversion rule that establishes a one-to-one correspondence between the planar orientation on the projected screen and the azimuth and elevation angles in the stereoscopic listening space. This is based on the fact that the visual field angle of the human eye and the auditory angle that the human ear can distinguish should be roughly aligned at the viewing distance, so that the left and right ends on the screen coincide with the left and right ends in the listening space.
[0055] In one embodiment, the preset value of the maximum recognizable azimuth angle on one side in the horizontal plane of the listener is denoted as α, which is 30 degrees. This value is approximately half of the horizontal angle of 52 degrees calculated based on a viewing distance of 3 meters and a projection screen width of 2.7 meters. The preset value of the maximum recognizable elevation angle on one side in the vertical plane of the listener is denoted as β, which is 15 degrees. This value is approximately half of the vertical angle of 28 degrees calculated based on a projection screen height of 1.5 meters. During the installation and debugging of the audio system, α and β are written into the system's configuration file using a configuration tool and remain unchanged during subsequent operation. The specific form of the audiovisual mapping relationship is as follows: the target azimuth angle θ = px × α, the sign of θ is inherited from px, the negative value points to the left of the listener and the positive value points to the right of the listener; the target elevation angle φ is obtained by subtracting the center offset of 0.5 from py and then multiplying by β. Since the value range of py is 0 to 1, where 0.5 corresponds to the vertical center position of the screen, when py equals 0.5, φ is zero, which is the horizontal plane of the ear. When py is greater than 0.5, φ is a positive value, indicating that the position of the mouth shape in the screen is above the horizontal plane of the listener's ear. When py is less than 0.5, φ is a negative value, indicating that it is below it.
[0056] It is understandable that when the position of the lip movements in the image falls exactly on the vertical center line of the projected image, px is zero, θ is zero, and the target sound image is located directly in front of the center line; when the position of the lip movements in the image falls exactly on the vertical center of the projected image, py is 0.5, φ is zero, and the target sound image is located on the horizontal plane of the listener's ear.
[0057] Preferably, in a three-dimensional listening space coordinate system with the preset listening position as the origin, the direction directly in front of the center line as the positive vertical axis, the direction pointing to the right of the listener in the horizontal plane as the positive horizontal axis, and the direction pointing vertically upward as the positive vertical axis, a unit vector pointing from the origin to the positive vertical axis is taken as the initial direction vector. The initial direction vector is first rotated in the horizontal plane around the vertical axis according to the target azimuth angle θ, and then raised or lowered in the vertical plane around the horizontal axis according to the target elevation angle φ, to obtain a spatial ray pointing from the origin to the target azimuth. Along the spatial ray, the distance corresponding to the perception radius R is traveled from the origin of the coordinate system along the ray direction. Here, R represents the preset perception radius used to characterize the estimated scale of the listener's subjective perception of the sound source distance. It is the same physical quantity as the perception radius used when the initial sound image landing point was marked, and the value is 1.5 meters. This yields a directional point located in the stereo listening space. The directional point is the directional point at which the sound should be perceived by the listener when the direction of the sound source matches the position of the lip movements in the image. This yields the target sound image landing point.
[0058] Extract the imaging coordinates of the person's face on the projected screen, identify the lateral offset of the mouth shape position from the vertical line of the center of the projected screen and the vertical height from the bottom edge of the projected screen, evaluate the degree of alignment of the mouth shape position on the projected screen to the left or right within the visible area of the projected screen, and obtain the planar orientation of the mouth shape position on the projected screen.
[0059] The horizontal and vertical pixel components are extracted from the facial imaging coordinates corresponding to the lip-sync position in the image. The horizontal pixel component is subtracted from the pixel coordinates of the center vertical line of the projected image to obtain the horizontal offset of the lip-sync position from the center vertical line. The vertical pixel component is measured upwards from the bottom edge of the projected image to obtain the vertical height of the lip-sync position from the bottom edge. The sign of the horizontal offset determines whether the lip-sync position is biased to the left or right side of the projected image. The absolute value of the horizontal offset is divided by half the width of the projected image to obtain the edge-fit degree with a positive or negative sign. The edge-fit degree and the vertical height are used together to obtain the planar orientation of the lip-sync position on the projected image.
[0060] In a home audio-visual entertainment scenario, the aforementioned facial imaging coordinates are defined by the center pixel of the facial detection box on the projected screen, with the horizontal pixel component denoted as u and the vertical pixel component denoted as v.
[0061] In one possible implementation, the center vertical line of the projected image is located at a horizontal pixel coordinate of W / 2, where W is the total number of pixels in the width of the projected image; the bottom edge of the projected image is located at a vertical pixel coordinate of 0, and the vertical pixel coordinates increase upwards to the total number of pixels H at the top edge of the image. The horizontal offset of the lip shape position from the center vertical line is taken as uW / 2, with a positive sign indicating it is located to the right of the center vertical line and a negative sign indicating it is located to the left; the vertical height of the lip shape position from the bottom edge is directly taken as the vertical pixel component v.
[0062] Specifically, to distinguish it from the lateral offset ratio symbol used in the aforementioned audiovisual mapping relationship, the degree of edge contact is denoted as qx in this paragraph. The value is determined by dividing the absolute value of the lateral offset from the lip-sync position to the center vertical line by half the projection screen width W / 2, while retaining the original positive or negative sign of the lateral offset; that is, qx falls within a closed interval from -1 to +1. The absolute value of qx reflects the degree of edge contact between the lip-sync position and the left or right edges of the projection screen. The closer the absolute value of qx is to 1, the closer the lip-sync position is to the side edge. A qx of 0 indicates that the lip-sync position falls exactly on the center vertical line. qx is completely consistent with the aforementioned lateral offset ratio px in terms of both value and symbol; they are different paragraph symbols representing the same physical quantity.
[0063] Preferably, to eliminate the inconsistency in dimensions between the horizontal and vertical components, the vertical height v is divided by the total number of pixels H in the projected image height to obtain the vertical height ratio qy. qy falls within a closed interval of 0 to 1, and both qy and qx are dimensionless ratios. The planar orientation point of the image lip-sync on the projected image is formed by the edge-sync degree qx and the vertical height ratio qy, which together constitute a binary tuple as the normalized coordinate representation of the planar orientation point on the projected image. qx carries the left-right horizontal orientation information, and qy carries the up-down vertical orientation information, thus obtaining the planar orientation point of the image lip-sync on the projected image.
[0064] Step S105: Evaluate whether the initial sound image landing point and the target sound image landing point are separated from each other in the stereo listening space, identify the playback segment where the sound direction is separated from the lip movements on the screen, and obtain the sound image separation time period.
[0065] The Euclidean distance between the initial sound image landing point and the target sound image landing point in the stereo listening space is calculated from their coordinates. If the Euclidean distance is greater than a preset separation threshold, the current sound frame is marked as a detached frame; otherwise, it is marked as a aligned frame. The preset separation threshold is in meters and is determined with reference to the listener's identifiable boundary of sound image directional offset. Based on the timestamps corresponding to the detached frames, multiple detached frames with a timestamp interval less than twice the video frame period are grouped into the same time period. The start point of the time period is the earliest detached frame timestamp within that time period, and the end point is the latest detached frame timestamp within that time period. The aligned frames are not included in the time period merging process, resulting in a playback segment where the sound direction is detached from the lip movements on the screen, i.e., the sound image detachment time period.
[0066] In one possible implementation, the spatial coordinates of the initial sound image landing point are denoted as P1, and its horizontal, vertical, and longitudinal components are denoted as x1, y1, and z1, respectively; the spatial coordinates of the target sound image landing point are denoted as P2, and its horizontal, vertical, and longitudinal components are denoted as x2, y2, and z2, respectively. The calculation rule for the Euclidean distance D is as follows: Where sqrt represents the arithmetic square root, and D is in meters. The physical meaning of D is the spatial geometric distance between P1 and P2 in the stereo listening space.
[0067] Specifically, a preset separation threshold is denoted as D0, and its value is determined based on the viewing distance L and the angular resolution of the human eye for sound-image offset. When the sound-image offset angle exceeds 3 degrees, the listener can clearly perceive the audiovisual separation, and D0 is calculated as L × tan3°. In a living room viewing scenario, the viewing distance is typically 3 to 5 meters, corresponding to a D0 value range of 0.16 to 0.26 meters, with a typical value of 0.2 meters. In a bedroom viewing scenario, the viewing distance is typically 2 to 3 meters, corresponding to a D0 value range of 0.10 to 0.16 meters, with a typical value of 0.13 meters. If D is greater than D0, the current sound frame is marked as a disconnected frame; if D is less than or equal to D0, the current sound frame is marked as a aligned frame. The disconnected frame corresponds to the moment when the sound direction and the lip movements on the screen show audiovisual separation, and the aligned frame corresponds to the moment when the sound direction and the lip movements on the screen are basically aligned. Furthermore, the detached frames are arranged in chronological order by timestamp, and the video frame period is denoted as T, which is equal to the reciprocal of the video frame rate. For example, when the frame rate is 30 frames per second, T = 0.033 seconds. The time-segment merging adopts a chain-like expansion rule, with a merging threshold set to 2T based on the following principle: the persistence of vision in the human eye is approximately 0.05 to 0.2 seconds. Two consecutive detached frames with an interval less than 2T can be perceived as belonging to the same detached time-segment. This threshold effectively merges discontinuous frames belonging to the same separation phenomenon while avoiding misjudging independent separation events as continuous. For two detached frames with adjacent timestamps, if the timestamp interval is less than 2T, the latter detached frame is merged into the detached time-segment containing the former. For the last detached frame already merged into a certain detached time-segment, the timestamp interval of its next detached frame is checked. If it is still less than 2T, it continues to be merged into that time-segment until a detached frame with an interval greater than or equal to 2T appears. At this point, a new detached time-segment is started. If increased separation detection sensitivity is required in practical applications, the merging threshold can be lowered to 1.5T; if a lower false alarm rate is required, it can be increased to 3T. The merging frames do not participate in the merging process during the separation period.
[0068] Preferably, for each separation period, the earliest separation frame timestamp in that period is taken as the start of the period, and the latest separation frame timestamp in that period is taken as the end of the period. The time interval between the start and end is the playback segment where the sound location separates from the lip movements on the screen, thus obtaining the sound-image separation period.
[0069] It is understandable that when the Euclidean distance D is continuously greater than D0 throughout the entire time period of a person speaking in the picture, all sound frames are detached frames, and after merging, a sound-image detached time period covering the entire speech is obtained; when the position of the mouth shape changes from left to right as the camera moves during the person's speech, detached frames and attached frames appear alternately, and after merging, several discontinuous sound-image detached time periods are obtained.
[0070] Step S106: For the audio signal within the time period of the sound image separation, based on the lateral placement of the first speaker and the second speaker, the initial sound image orientation, and the target sound image orientation, determine the channel time difference adjustment amount required to migrate the sound image orientation from the initial sound image orientation to the target sound image orientation, and use this as the second time difference.
[0071] For the audio signal within the time interval of the sound image separation, the target azimuth angle of the synthesized sound image relative to the center line of the preset listening position is extracted from the target sound image landing point. Using the sine correspondence between the binaural time difference and the azimuth angle, the target binaural time difference corresponding to the target azimuth angle is calculated back based on empirical values of binaural distance and sound velocity. The initial azimuth angle of the synthesized sound image relative to the center line is determined based on the initial sound image landing point. The initial binaural time difference corresponding to the initial azimuth angle is calculated back using the same sine correspondence. The initial binaural time difference is numerically consistent with the first time difference, reflecting the original timing difference of the channel signals between the first and second speakers in the current lateral placement position. Based on the value obtained by subtracting the initial binaural time difference from the target binaural time difference, and combined with the signal arrival order resulting from the lateral placement positions of the first and second speakers under the current listening geometry, the channel time difference adjustment amount required to migrate the sound image azimuth from the initial sound image landing point to the target sound image landing point is analyzed and used as the second time difference.
[0072] In home audio-visual entertainment scenarios, the audio signal during the time period when the sound image is separated consists of the left channel signal and the right channel signal. Next, the channel time difference adjustment required to migrate the sound image location from the initial sound image landing point to the target sound image landing point is calculated for this audio signal, and this is used as the input basis for subsequent delay processing.
[0073] In one possible implementation, the target azimuth angle θ2 is extracted from the spatial coordinates of the target sound image's landing point. θ2 is the angle between the projection of the line connecting the target sound image's landing point to the origin on the horizontal plane and the vertical axis. θ2 has a positive and a negative sign; a positive value points to the listener's right, and a negative value points to the listener's left. The absolute value of the target azimuth angle θ2 determines the horizontal deviation of the target sound image from the center line, and the sign determines whether the sound image should be biased to the left or right.
[0074] It should be noted that the sinusoidal correspondence between the binaural time difference and the azimuth angle is given by the classic spherical head model: when a sound source is incident on the listener at an azimuth angle θ, the time difference between the ear closer to the sound source and the other ear receiving the same sound wave is proportional to sinθ, and the proportionality coefficient is determined by empirical values of the binaural distance and the speed of sound. The binaural distance is denoted as d, which is taken as the average straight-line distance between the openings of the external auditory canals of the two ears in an adult, 0.175 meters; the empirical value of the speed of sound is denoted as c, which is taken as 340 meters per second.
[0075] Specifically, by substituting the target azimuth angle θ2 into the sine correspondence, the target binaural time difference t2 is obtained by reverse calculation, and the calculation rule is as follows: Where d is the head feature size, taken as 0.175 meters, c is the speed of sound, 340 meters per second, and the sign of t2 is inherited from the sign of θ2. The physical meaning of t2 is the time difference between when the listener perceives the sound source's location as falling on the target sound image's landing point, and when the left and right ears should receive the same sound wave. For example, when θ2 is positive 46 degrees, sinθ2 is approximately 0.719, and t2 = 0.175 * 0.719 / 340 ≈ 0.00037 seconds, or 0.37 milliseconds, which is exactly in the same listening geometry as the aforementioned first time difference example.
[0076] In one embodiment, an initial azimuth angle θ1 is extracted from the spatial coordinates of the initial sound image landing point. θ1 is defined using the same reference frame as the target azimuth angle θ2. The initial binaural time difference t1 is obtained by substituting the initial azimuth angle θ1 into the same sine correspondence, calculated as t1 = d*sinθ1 / c. The initial binaural time difference t1 describes the time difference at which the left and right ears should receive the same sound wave when the listener's actual perceived sound source location falls on the initial sound image landing point.
[0077] It should be noted that t1 is not recalculated, but rather the first time difference value obtained above is directly used. The premise that the binaural time difference t1 is numerically equivalent to the initial inter-channel time difference is that the preset listening position is located near the perpendicular bisector of the line connecting the first and second speakers. In this case, the time difference between the left and right channel signals reaching the listening position can be equivalent to the time difference between the same synthesized sound source reaching the listener's ears. The two share the same proportional coefficient through the sinusoidal relationship d×sinθ / c mentioned above, avoiding the consistency risk caused by the same physical quantity being obtained independently through two different paths in the scheme. Furthermore, the inter-channel time difference adjustment Δt is defined as the additional timing difference introduced between the left and right channel signals required to migrate the listener's perceived orientation from the initial sound image landing point to the target sound image landing point. The target sound image landing point refers to the sound source perception position that the user expects to adjust to, and its azimuth angle θ2 is set by the user through the interactive interface or automatically assigned by the system according to the content type. The formula is used when calculating the binaural time difference t2. The calculation rule is Δt = t2 - t1. The sign of Δt determines whether the subsequent delay processing is applied to the left channel signal or the right channel signal. If Δt is positive, it means that the left channel signal should arrive at the listener's left ear with an additional delay relative to the right channel signal. That is, the right channel signal is considered to arrive first, to simulate the time difference between the two ears when the sound source is on the right side of the listener. If Δt is negative, it means that the right channel signal should arrive at the listener's right ear with an additional delay relative to the left channel signal. That is, the left channel signal is considered to arrive first, to simulate the time difference between the two ears when the sound source is on the left side of the listener.
[0078] Specifically, the lateral placement of the first and second speakers determines the arrival time difference of the channel signals during free space propagation, i.e., the aforementioned first time difference. The channel time difference adjustment Δt acts on the electrical signal level, pre-adjusting the timing of the channel signals entering the power amplifier channel before the speakers emit sound waves. After the spatial propagation time difference and the electrical signal time difference are linearly superimposed at the preset listening position, the total time difference reaching the listener's ears should be equal to the target binaural time difference t2. That is, the sum of Δt and the first time difference should be numerically equal to t2. This equation is the original source of Δt = t2 - t1.
[0079] Preferably, considering that the human ear's ability to distinguish time differences between audio channels is approximately 0.01 milliseconds, Δt is rounded to the nearest multiple of 0.01 milliseconds as the final second time difference. For example, when t2 is 0.37 milliseconds and t1 is 0.12 milliseconds, Δt = 0.37 - 0.12 = 0.25 milliseconds, and the second time difference is 0.25 milliseconds. A positive value indicates that the right channel signal should be delayed further than the left channel signal.
[0080] It is understandable that when the initial sound image landing point coincides with the target sound image landing point, θ2 is equal to θ1, t2 is equal to t1, and the second time difference is zero. At this time, there is no need to introduce additional inter-channel time difference adjustment. When the target sound image landing point is on the same side as the initial sound image landing point but closer to the midline, the sign of the second time difference is opposite to t1, which is used to pull the deviated sound image back to the center. When the target sound image landing point is on the opposite side of the initial sound image landing point, the absolute value of the second time difference is significantly greater than the absolute value of t1, which is used to migrate the sound image across the midline to the opposite side.
[0081] In one embodiment, when a person in the picture moves from the left side to the right side of the projected picture as the camera switches, the target azimuth angle θ2 changes from a negative value to a positive value, and the sign of the second time difference is also flipped, which is used to pull the sound image from the left side of the picture to the right side of the picture, and the adjustment amount of the time difference between the channels during the current sound image separation time period is used as the second time difference.
[0082] Step S107: Based on the second time difference, apply a corresponding delay processing to the channel signal that arrives at the preset listening position first in the audio signal, so that the time difference between the left and right channel signals arriving at the preset listening position after processing matches the target sound image location, and obtain an output audio signal that re-matches the sound location with the lip shape on the screen and sends it to the first speaker and the second speaker.
[0083] Based on the sign of the second time difference, the channel signal that arrives at the preset listening position first is determined from the left and right channel signals of the audio signal. If the second time difference is positive, the left channel signal is determined to be the first arriving channel; if the second time difference is negative, the right channel signal is determined to be the first arriving channel. The number of samples to be delayed is obtained by multiplying the absolute value of the second time difference by the sampling rate of the audio signal and rounding down. The sampling rate is measured in samples per second. A digital delay buffer is used to apply a sample delay consistent with the number of samples to the first arriving channel, while the other channel signal maintains its original timing, resulting in delayed left and right channel signals. The left channel signal from the delayed left and right channel signals is sent to the amplifier channel corresponding to the first speaker, and the right channel signal is sent to the amplifier channel corresponding to the second speaker, resulting in an output audio signal whose sound direction is re-matched to the lip movements on the screen. This output audio signal is then played through the first and second speakers.
[0084] Next, the audio signal that arrives at the preset listening position first in the audio signal during the time period of sound image separation is subjected to corresponding delay processing, so that the time difference between the left and right channel signals arriving at the preset listening position after processing matches the target sound image location, and the output audio signal that re-matches the sound location with the lip shape of the picture is obtained and sent to the first speaker and the second speaker for playback.
[0085] In one possible implementation, the selection of the first-arriving channel is determined based on the sign of the second time difference. When the second time difference is positive, the right channel signal should arrive at the listener's right ear with an additional delay relative to the left channel signal; that is, the left channel signal is considered to arrive first, and in this case, the left channel signal is taken as the first-arriving channel. When the second time difference is negative, the right channel signal is taken as the first-arriving channel. When the second time difference is zero, there is no need to select a first-arriving channel, and the delay application process is skipped directly.
[0086] It should be noted that the implementation of sample delay relies on a digital delay buffer. A digital delay buffer is a common type of sample buffer in the field of digital audio processing that operates on a first-in, first-out (FIFO) rule. The sample points written first at the writing end are read first at the reading end. The buffer length can be dynamically configured. The writing end receives the original sample points of the channel signal, and the reading end outputs the lagging sample points according to the number of lagging samples between the writing end and the reading end. The sample delay is the number of sample intervals between the writing end and the reading end.
[0087] Specifically, the number of samples N is obtained by multiplying the absolute value of the second time difference t by the sampling rate fs of the audio signal and rounding down, i.e., N = floor(tfs), where floor represents the rounding down operation, t is in seconds, fs is in the number of samples per second, the product tfs is in the number of samples, and N is the integer number of samples obtained by rounding down.
[0088] For example, when the second time difference is 0.25 milliseconds (t = 0.00025 seconds) and the sampling rate is 48,000 samples per second (fs = 48,000), then t * fs = 12, and N is 12 sampling points. When the sampling rate is 96,000 samples per second, N corresponds to 24 sampling points for the same time difference. The number of samples increases linearly with the increase of the sampling rate, and the corresponding time delay accuracy also improves.
[0089] In one embodiment, the original sampling point sequence of the first-arriving channel is sequentially written to the write end of the digital delay buffer, and the read end retrieves sampling points from the buffer according to the read pointer lagging N sampling points as the delayed first-arriving channel signal; the other channel signal is directly transmitted without going through the digital delay buffer, that is, the other channel signal maintains its original timing. The two channel signals are re-aligned in time and then merged into the delayed left and right channel signals.
[0090] Preferably, to avoid transient noise caused by amplitude jumps in the channel signal during delay parameter switching, a linear transition is used when the number of samples N switches from the current value to the new value. The preset transition sample number M is determined based on the human hearing masking threshold, and its value ranges from 0.5 to 2 milliseconds of the sampling rate to the number of sampling points. Under the condition of a sampling rate of 48,000 sampling points per second, M is taken as 64 sampling points, corresponding to a smoothing duration of 1.33 milliseconds. This duration is less than the human ear's time resolution threshold of 3 milliseconds for transient signals. The read pointer indicates the current position index of the sample read from the circular buffer. Before the switch, the read pointer and write pointer are separated by N1 old sample points to form the old lag position. After the switch, they are separated by N2 new sample points to form the new lag position. Within the transition interval of M sample points, the read pointer position moves according to the linear interpolation rule L = L1 + k * L2 - L1 / M, where k is the current transition progress increasing from 0 to M - 1, L1 is the read pointer index value corresponding to the old lag position, and L2 is the read pointer index value corresponding to the new lag position. Furthermore, the channel allocation maintains the original spatial positions of the left and right channels. After the delay, the left channel signal is amplified by the left power amplifier channel and sent to the electroacoustic transducer unit of the first speaker, and the right channel signal is amplified by the right power amplifier channel and sent to the electroacoustic transducer unit of the second speaker. The first and second speakers are fixed in the living room with a horizontal spacing of 2 to 5 meters, and no physical position adjustment is involved during playback.
[0091] It is understandable that the two audio signals will generate an initial timing difference, i.e., a first time difference, during free space propagation, determined by the lateral placement position. By applying a sample-level delay to one of the signals, a second time difference is introduced at the electrical signal level. This sample delay is achieved through a digital delay buffer, specifically by converting the delay duration into an integer number of samples according to the sampling rate before buffering and storing / reading. The first and second time differences are linearly superimposed at the preset listening position, and the resulting total time difference reaching the listener's ears is numerically equal to the target binaural time difference. The target binaural time difference is determined by the desired sound source azimuth angle, which is preset according to the sound field reproduction requirements. For example, if the desired virtual sound source is located 30 degrees to the left of the listener, the corresponding target binaural time difference is approximately 0.26 milliseconds. This target value is achieved by superimposing the first and second time differences, so the sound source location perceived by the listener coincides with the desired location.
[0092] In one embodiment, a multimedia signal including an audio stream and a video stream is received. A face detection algorithm in the video stream is used to identify people in the scene, and a lip-sync detection algorithm is used to determine the person's speaking state. When a lip-sync movement amplitude exceeds a preset threshold of 0.3 and its duration exceeds 100 milliseconds, it is determined that the person has begun speaking and enters the audio-visual separation period. During this period, the left or right channel signal is delayed using a digital delay buffer. The digital delay buffer refers to a variable-length sample buffer queue set in the audio signal processing path, whose lag sample count is updated in real time according to the current second time difference. When a lip-sync movement amplitude is detected to be less than 0.3 or the speaking duration exceeds the expected end time of the current sentence, it is determined that the speaking has ended or the separation period has ended. At this time, the lag sample count smoothly drops back to 0 according to the aforementioned transition sample count M, meaning that the left and right channel signals return to a state of error-free delayed transmission. The above processing is performed in the main audio rendering loop, which runs periodically at the time interval corresponding to the audio sampling rate. For example, it executes 48,000 times per second when the sampling rate is 48,000 Hz. Audio and video synchronization is achieved by reading the timestamps of the video stream and the audio stream. Delay processing is triggered and released according to the start and end timestamps of the time period when the sound and image are separated. The output audio signal is obtained by re-matching the sound direction with the lip movements on the screen and played through the first speaker and the second speaker.
[0093] This invention provides a split-type sound projection system, mainly comprising: The first acquisition module is used to acquire the horizontal placement positions of the first speaker located on one side of the center line of the screen and the second speaker located on the other side of the center line of the screen relative to the center line of the screen, and to calculate the first time difference between the left channel signal and the right channel signal reaching the preset listening position. The first calculation module is used to identify the sound source location of the left and right channel signals based on the first time difference, and determine the initial sound image landing point of the synthesized left and right channel signals in the stereo listening space. The lip-reading recognition module is used to recognize the lip movements of people in a video frame that is played synchronously with the left and right channel signals, and to identify the position of the lip movements of the person who is currently speaking. The target sound image determination module is used to obtain the planar orientation of the lip shape position on the projected screen, and determine the target sound image landing point where the lip shape should be uttered based on the planar orientation and the preset audiovisual mapping relationship. The evaluation module is used to evaluate whether the initial sound image landing point and the target sound image landing point are separated from each other in the stereo listening space, and to obtain the sound image separation time period. The positioning module is used to locate the time difference adjustment between the sound channels for the audio signal within the time period of the sound image separation, based on the horizontal placement position, the initial sound image orientation and the target sound image orientation, as the second time difference; The delay processing module is used to apply delay processing to the channel signal that arrives at the preset listening position first in the audio signal according to the second time difference, so as to obtain the output audio signal and send it to the first speaker and the second speaker.
[0094] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A split-type sound projection method, characterized in that, The method includes: Obtain the lateral placement of the first speaker located on one side of the center line of the screen and the second speaker located on the other side of the center line of the screen relative to the center line of the screen, and calculate the first time difference between the left channel signal and the right channel signal reaching the preset listening position. Based on the first time difference, the sound source location of the left and right channel signals is identified to determine the initial sound image landing point in the stereo listening space after the left and right channel signals are synthesized. Lip-reading recognition is performed on the faces of people in video images that are played synchronously with the left and right channel signals to identify the position of the lip movements of the person currently speaking. The planar orientation of the lip movements in the image is obtained on the projected image. Based on the planar orientation and the preset audiovisual mapping relationship, the target sound image landing point where the lip movements in the image should be pronounced is determined. Evaluate whether the initial sound image landing point and the target sound image landing point are separated from each other in the stereo listening space, and obtain the sound image separation time period; For the audio signal within the time period of the separation of the sound image, the time difference adjustment amount between the channels is determined based on the horizontal placement position, the initial sound image orientation and the target sound image orientation, and is used as the second time difference; Based on the second time difference, a delay is applied to the channel signal that arrives at the preset listening position first in the audio signal to obtain the output audio signal and send it to the first speaker and the second speaker.
2. The split-type sound projection method according to claim 1, characterized in that, The process of obtaining the lateral placement positions of the first speaker located on one side of the center line of the image and the second speaker located on the other side of the center line relative to the center line of the image, and calculating the first time difference between the left channel signal and the right channel signal reaching the preset listening position, includes: deploying infrared ranging sensors from the vertical line of the center of the projected image to both sides to obtain the lateral placement positions of the first speaker and the second speaker relative to the vertical line of the center; forming a right-angled triangle in the horizontal plane with the lateral placement positions and the longitudinal distance from the preset listening position to the vertical line of the center, and using the Pythagorean theorem to calculate the path lengths from the first speaker and the second speaker to the preset listening position; and dividing the difference between the two path lengths from the first speaker and the second speaker to the preset listening position and the second speaker to the preset listening position by the speed of sound to obtain the first time difference.
3. The split-type sound projection method according to claim 1, characterized in that, The step of identifying the sound source location of the left and right channel signals based on the first time difference and determining the initial sound image landing point in the stereo listening space after the left and right channel signals are synthesized includes: determining the offset direction of the synthesized sound image relative to the center line of the preset listening position based on the sign of the first time difference; establishing a horizontal coordinate system of the stereo listening space with the preset listening position as the origin and the center line as the vertical axis; calculating the azimuth angle of the synthesized sound image relative to the vertical axis using an arcsine function based on the ratio of the absolute value of the first time difference to the interaural distance; and marking the azimuth point along the vertical axis with a preset sensing radius as the distance and the azimuth angle as the subtended angle in the offset direction to obtain the initial sound image landing point.
4. The split-type sound projection method according to claim 1, characterized in that, The step of performing lip-syncing on a person's face in a video frame played synchronously with the left and right channel signals to identify the lip position of the person currently speaking includes: extracting a video frame sequence frame by frame from the video frame played synchronously with the left and right channel signals; performing face detection on each frame using a multi-task cascaded convolutional neural network to obtain a face detection box; for the face image of the person within the face detection box, using a face key point localization method based on a depth alignment network to mark the coordinates of key points on the upper lip contour, lower lip contour, and corners of the mouth; determining the opening and closing amplitude of the lip region by the vertical pixel distance between the central key point of the upper lip and the central key point of the lower lip; if the difference in opening and closing amplitude between adjacent video frames is greater than a preset sound threshold, it is determined as a sounding frame and retained; otherwise, it is determined as a silent frame and discarded; for the retained sounding frames, the centroid of the region enclosed by the central key points of the upper lip, central key points of the lower lip, and corners of the mouth is used as the lip coordinates, and aligned synchronously with the left and right channel signals according to the video frame timestamp to obtain the lip position of the image.
5. The split-type sound projection method according to claim 1, characterized in that, The step of obtaining the planar orientation of the lip-sync position on the projected screen, and determining the target sound image landing point where the lip-sync should be produced based on the planar orientation and a preset audiovisual mapping relationship, includes: obtaining a lateral offset ratio with positive and negative signs from the offset of the lateral pixel component relative to the center vertical line, and obtaining a vertical height ratio from the height of the vertical pixel component relative to the bottom edge, as the planar orientation; using the preset audiovisual mapping relationship, converting the lateral offset ratio into the target azimuth angle, and converting the vertical height ratio into the target elevation angle; and marking the target sound image landing point along a preset perception radius direction according to the target azimuth angle and the target elevation angle, with the preset listening position as the origin and the direction directly in front of the center line as the reference.
6. The split-type sound projection method according to claim 1, characterized in that, The method further includes: evaluating the degree to which the lip-sync point on the projected image is positioned to the left or right within the visible area of the projected image, thereby obtaining the planar orientation of the lip-sync point on the projected image.
7. The split-type sound projection method according to claim 1, characterized in that, The step of evaluating whether the initial sound image landing point and the target sound image landing point are separated from each other in the stereo listening space to obtain the sound image separation time period includes: calculating the Euclidean distance between the initial sound image landing point and the target sound image landing point in the stereo listening space based on the coordinates of the initial sound image landing point and the target sound image landing point; marking the current sound frame as a separation frame when the Euclidean distance is greater than a preset separation threshold, and marking it as a fit frame otherwise; according to the timestamps corresponding to the separation frames, grouping multiple separation frames with an interval of less than twice the video frame period into the same time period, taking the earliest separation frame timestamp as the start point of the time period and the latest separation frame timestamp as the end point of the time period; and not including the fit frames in the time period merging to obtain the sound image separation time period.
8. The split-type sound projection method according to claim 1, characterized in that, The method for determining the inter-channel time difference adjustment amount as the second time difference for audio signals within the time period of sound image separation, based on the lateral placement position, initial sound image orientation, and target sound image orientation, includes: extracting the target azimuth angle of the synthesized sound image relative to the centerline from the target sound image landing point; using the sine correspondence between the binaural time difference and the azimuth angle, and back-calculating the target binaural time difference corresponding to the target azimuth angle based on the binaural distance and sound velocity values; determining the initial azimuth angle from the initial sound image landing point, and back-calculating the initial binaural time difference using the same sine correspondence, wherein the initial binaural time difference is consistent with the first time difference; subtracting the initial binaural time difference from the target binaural time difference, and combining this with the signal arrival order caused by the lateral placement position, to obtain the second time difference.
9. The split-type sound projection method according to claim 1, characterized in that, The step of applying a delay to the audio signal that arrives at the preset listening position first based on the second time difference to obtain an output audio signal and sending it to the first and second speakers includes: determining the first channel to arrive at the preset listening position from the left and right channel signals based on the sign of the second time difference; multiplying the absolute value of the second time difference by the sampling rate of the audio signal and rounding down to obtain the number of delay samples; applying a sample delay consistent with the number of samples to the first channel using a digital delay buffer; maintaining the original timing of the other channel signal; and after the delay, sending the left channel signal to the first speaker amplifier channel and the right channel signal to the second speaker amplifier channel to obtain the output audio signal.
10. A split-type audio projection system, characterized in that, The system includes: The first acquisition module is used to acquire the horizontal placement positions of the first speaker located on one side of the center line of the screen and the second speaker located on the other side of the center line of the screen relative to the center line of the screen, and to calculate the first time difference between the left channel signal and the right channel signal reaching the preset listening position. The first calculation module is used to identify the sound source location of the left and right channel signals according to the first time difference, and determine the initial sound image landing point of the synthesized left and right channel signals in the stereo listening space. The lip-reading recognition module is used to recognize the lip movements of people in a video frame that is played synchronously with the left and right channel signals, and to identify the position of the lip movements of the person who is currently speaking. The target sound image determination module is used to obtain the planar orientation of the lip shape position on the projected screen, and determine the target sound image landing point where the lip shape should be uttered based on the planar orientation and the preset audiovisual mapping relationship. The evaluation module is used to evaluate whether the initial sound image landing point and the target sound image landing point are separated from each other in the stereo listening space, and to obtain the sound image separation time period. The positioning module is used to locate the time difference adjustment between the sound channels for the audio signal within the time period of the sound image separation, based on the horizontal placement position, the initial sound image orientation and the target sound image orientation, as the second time difference; The delay processing module is used to apply delay processing to the channel signal that arrives at the preset listening position first in the audio signal according to the second time difference, so as to obtain the output audio signal and send it to the first speaker and the second speaker.