A method, apparatus, device and medium for adjusting a sound source based on audience position
By determining the location information of the sound source image and the audience, calculating the intersection of the lines and controlling the audio beam, the problem of audio-visual asynchrony was solved, achieving audio-visual synchronization and the best auditory experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA TELECOM CLOUD TECH CO LTD
- Filing Date
- 2024-12-02
- Publication Date
- 2026-05-29
AI Technical Summary
When viewers move away from the dessert area of the screen, the sound field of the speakers deteriorates sharply, causing audio and video to become out of sync and affecting the viewing experience.
By determining the location information of the sound source image and the audience, the intersection point of their line is calculated, and multiple target ultrasonic transmitting units are selected based on the intersection point to control the direction and width of the audio beam, so as to achieve precise orientation of the audio output.
It achieves audio-visual synchronization, enhances the audience's auditory experience, and ensures that every audience member receives the best auditory effect.
Smart Images

Figure CN119729333B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound source adjustment technology based on audience position, and in particular to a sound source adjustment method, apparatus, device, and medium based on audience position. Background Technology
[0002] With the development of technology, the types and functions of audio-visual products are increasing, and viewers' requirements for the effects of audio-visual products are also increasing. However, when viewers deviate from the sweet spot of the screen, the sound field produced by the speakers deteriorates sharply, creating a clear sense of spatial separation between the sound and the displayed image, resulting in audio-visual asynchrony and affecting the viewing experience. Summary of the Invention
[0003] In view of the above problems, embodiments of the present invention are proposed to provide a method, apparatus, device and medium for adjusting a sound source based on the audience position to overcome the above problems or at least partially solve the above problems.
[0004] To address the aforementioned problems, embodiments of the present invention disclose a sound source adjustment method based on audience position, comprising:
[0005] During the process of displaying images on the screen, determine the location information of the sound source image and the location information of the audience;
[0006] Determine the location information of the sound source image and the location information of the audience, and extend the line connecting them to the intersection on the screen;
[0007] Based on the intersection point, a plurality of target ultrasonic transmitting units are determined from a plurality of ultrasonic transmitting units arranged along the screen;
[0008] Based on the location information of the audience, the audio beams output by the multiple target ultrasonic transmitting units are controlled respectively.
[0009] Optionally, determining the location information of the sound source image includes:
[0010] Determine the type of sound source;
[0011] When the sound source type is human, determine the center point position of the mouth of the human face in the sound image, and use the center point position of the mouth of the human face as the position information of the sound source image;
[0012] When the sound source type is non-human, the horizontal center point position of the sound image is determined, and the horizontal center point position is used as the position information of the sound source image.
[0013] Optionally, the plurality of ultrasonic emitting units are arranged along the lateral edge of the screen; the step of determining a plurality of target ultrasonic emitting units from the plurality of ultrasonic emitting units arranged along the screen based on the intersection point includes:
[0014] Along the vertical direction of the intersection point, a target area is determined within the deployment area of the plurality of ultrasonic transmitting units;
[0015] The ultrasonic transmitting unit located within the target area is identified as the target ultrasonic transmitting unit.
[0016] Optionally, controlling the audio beams output by the plurality of target ultrasonic transmitting units according to the location information of the audience includes:
[0017] Based on the location information of the audience, the direction of the audio beams output by the plurality of target ultrasonic transmitting units is controlled to be towards the audience.
[0018] Optionally, the step of controlling the audio beams output by the plurality of target ultrasonic transmitting units according to the location information of the audience further includes:
[0019] Based on the location information of the audience, the distances of the plurality of target ultrasonic transmitting units relative to the audience are determined;
[0020] The width of the audio beams output by the multiple target ultrasonic transmitting units is controlled according to the distance between the multiple target ultrasonic transmitting units and the audience.
[0021] Optionally, the method further includes:
[0022] When the position information of the audience remains unchanged, but the position information of the sound source image changes, the angle of change in the direction of the audio beams output by the multiple target ultrasonic transmitting units is detected.
[0023] When the change angle is greater than a preset angle threshold, the target ultrasonic transmitting unit is redefined.
[0024] Optionally, the method further includes:
[0025] When the location information of the audience changes, but the location information of the sound source image remains unchanged, determine whether the location information of the audience exceeds the coverage range of the audio beams output by the multiple target ultrasonic transmitting units;
[0026] When the location information of the audience exceeds the coverage range of the audio beams output by the plurality of target ultrasonic transmitting units, the target ultrasonic transmitting units are re-determined.
[0027] Accordingly, embodiments of the present invention also disclose a sound source adjustment device based on audience position, comprising:
[0028] The position information determination module is used to determine the position information of the sound source image and the position information of the audience during the process of displaying images on the screen.
[0029] The intersection point determination module is used to determine the intersection point of the line connecting the position information of the sound source image and the position information of the audience, extending to the screen.
[0030] A target ultrasonic transmitting unit determination module is used to determine multiple target ultrasonic transmitting units from multiple ultrasonic transmitting units arranged along the screen based on the intersection point;
[0031] The audio beam control module is used to control the audio beams output by the multiple target ultrasonic transmitting units according to the position information of the audience.
[0032] Optionally, the intersection point determination module includes:
[0033] The sound source type determination submodule is used to determine the sound source type;
[0034] The first location information determination submodule is used to determine the center point position of the mouth of a human face in the sound image when the sound source type is human, and to use the center point position of the mouth of the human face as the location information of the sound source image.
[0035] The second location information determination submodule is used to determine the horizontal center point position of the sound image when the sound source type is non-human, and use the horizontal center point position as the location information of the sound source image.
[0036] Optionally, the plurality of ultrasonic emitting units are arranged along the lateral edge of the screen; the target ultrasonic emitting unit determining module includes:
[0037] The target area determination submodule is used to determine the target area along the vertical direction of the intersection point within the deployment area of the plurality of ultrasonic transmitting units;
[0038] The target ultrasonic transmitting unit submodule is used to identify ultrasonic transmitting units located within the target area as target ultrasonic transmitting units.
[0039] Optionally, the audio beam control module includes:
[0040] The audio beam direction control submodule is used to control the direction of the audio beams output by the plurality of target ultrasonic transmitting units toward the audience, based on the audience's position information.
[0041] Optionally, the audio beam control module further includes:
[0042] The distance determination submodule is used to determine the distance between the plurality of target ultrasonic transmitting units and the audience based on the audience's location information;
[0043] The width determination submodule is used to control the width of the audio beams output by the multiple target ultrasonic transmitting units based on the distances of the multiple target ultrasonic transmitting units relative to the audience.
[0044] Optionally, the method further includes:
[0045] An angle detection module is used to detect the change angle of the direction of the audio beams output by the multiple target ultrasonic transmitting units when the position information of the audience remains unchanged and the position information of the sound source image changes.
[0046] The first re-determination submodule is used to re-determine the target ultrasonic transmitting unit when the changed angle is greater than a preset angle threshold.
[0047] Optionally, the method further includes:
[0048] The coverage area determination module is used to determine whether the location information of the audience exceeds the coverage area of the audio beams output by the multiple target ultrasonic transmitting units when the location information of the audience changes but the location information of the sound source image remains unchanged.
[0049] The second re-determination submodule is used to re-determine the target ultrasonic transmitting unit when the location information of the audience exceeds the coverage range of the audio beams output by the plurality of target ultrasonic transmitting units.
[0050] The embodiments of the present invention have the following advantages:
[0051] In this embodiment of the invention, firstly, the positional information of the sound source image and the positional information of the audience in the screen display are acquired; then, by analyzing this positional information, the line connecting the positional information of the sound source image and the positional information of the audience is determined and extended to the intersection point on the screen; then, based on the identified intersection point, multiple target ultrasonic transmitting units are determined from multiple ultrasonic transmitting units arranged along the screen; finally, based on the audience's positional information, the audio beams output by the multiple target ultrasonic transmitting units are controlled respectively. In this way, this embodiment of the invention can comprehensively, quickly, and accurately identify and process the relationship between the sound source image and the audience's position, enhancing the intelligence of audio output, achieving audio-visual synchronization, and ensuring that the audience obtains the best auditory experience. Attached Figure Description
[0052] Figure 1This is a flowchart illustrating the steps of a sound source adjustment method based on audience position provided in an embodiment of the present invention;
[0053] Figure 2 This is a flowchart of another sound source adjustment method based on audience position provided in an embodiment of the present invention;
[0054] Figure 3 This is a flowchart illustrating the specific steps of a sound source adjustment method based on audience position provided in an embodiment of the present invention.
[0055] Figure 4 This is a flowchart illustrating the specific steps of another sound source adjustment method based on audience position provided in an embodiment of the present invention.
[0056] Figure 5 This is a structural block diagram of a sound source adjustment device based on audience position provided in an embodiment of the present invention. Detailed Implementation
[0057] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0058] With the development of technology, the types and functions of audio-visual products are increasing, and viewers' demands for the effects of these products are also rising. However, when viewers move away from the sweet spot of the screen, the sound field produced by the speakers deteriorates sharply, creating a noticeable spatial separation between the sound and the displayed image, resulting in audio-visual asynchrony and affecting the viewing experience.
[0059] Traditional solutions are proving inadequate, as these problems involve not only the expansion of audio-visual equipment but also the complexity of sound field management. Traditional audio-visual systems have failed to fully utilize information and digital technologies to monitor and optimize sound field effects in real time, thus failing to effectively enhance the audience's auditory experience and causing inconvenience to people's daily lives.
[0060] One of the core concepts of this invention lies in its ability to acquire, in real time, the location information of the sound source image on the screen, as well as the location information of the audience. Simultaneously, it can accurately identify and process abnormal areas of the sound field by controlling speakers or linking with other devices. By precisely analyzing this location information, the relationship between the sound source image and the audience's position can be accurately identified. Based on these identification results, the audio-visual system will display corresponding processing strategies for rapid response. In this way, abnormal sound field conditions can be quickly and accurately identified and processed, intelligently achieving audio-visual synchronization and ensuring the audience receives the best auditory experience.
[0061] Reference Figure 1The diagram illustrates a flowchart of a sound source adjustment method based on audience position according to an embodiment of the present invention. The method may specifically include the following steps:
[0062] Step 101: During the process of displaying the image on the screen, determine the location information of the sound source image and the location information of the audience.
[0063] The method of this invention can be applied to intelligent audio-visual systems or terminal devices. The screen of the intelligent audio-visual system or terminal device can be fixed on the wall, and the display screen used to play the picture can be a flat, curved, or naked-eye 3D screen.
[0064] In some examples, the location of the audience can be determined using optical (camera), ultrasonic, electromagnetic fields, or other methods.
[0065] In some examples, for a flat-panel display, the lines connecting viewers at different positions to the sound source image point to the same intersection on the screen, so the intersection position is unique for multiple viewers in a flat-panel display scenario; for sound source images in front of / behind the screen (i.e., glasses-free 3D displays), the lines connecting viewers at different positions to the sound source image are not the same.
[0066] In some instances, the location of the sound source image corresponds to the audience's psychological expectation of the sound source. The human auditory system has the highest accuracy in locating sound sources in the horizontal direction in front, and the accuracy decreases rapidly from the center to the sides; however, it is not sensitive to the vertical movement of the sound source. Sometimes, the sound source movement will only be perceived after a 60° angular displacement at the subject's location. For screens with a vertical viewing angle much smaller than 60°, it is only necessary to consider the horizontal sound-image co-position, and the sound-image difference caused by different vertical sound source image positions can be ignored.
[0067] The method of this invention can use parametric array loudspeakers at different positions to emit sound in a directional manner for different audiences, and it works best for non-planar displays such as naked-eye 3D displays.
[0068] Step 102: Determine the line connecting the location information of the sound source image and the location information of the audience, extending it to the intersection point on the screen;
[0069] In this embodiment of the invention, based on the location information of the sound source image and the location information of the audience, a series of calculations and analyses can be performed to determine the line connecting the two and extend it to the intersection point on the screen. This step is a key step in achieving accurate sound source localization and audience interaction.
[0070] In some examples, the location of the sound source image and the audience's location can first be obtained through step 101. Next, the location information of the sound source image and the audience's location information can be matched and calculated. Using geometric principles, a line connecting the sound source and the audience can be determined and extended to its intersection on the screen. This intersection represents the projected location of the sound source on the screen, thus achieving a visual representation of the sound source.
[0071] In some examples, the line connecting the sound source and the audience can be calculated using geometric calculations and projection techniques: first, the three-dimensional coordinates of the sound source and the audience are calculated using sensor data; then, the equation of the line connecting the sound source and the audience is calculated using geometric principles; finally, the line is extended to the intersection on the screen to calculate the projection position.
[0072] In some examples, environmental factors, such as the sound's propagation path and reflections, can also be considered to further improve the accuracy of sound localization. These steps enable precise interaction between the sound source and the audience, providing a more immersive experience.
[0073] Step 103: Based on the intersection point, determine a plurality of target ultrasonic transmitting units from a plurality of ultrasonic transmitting units arranged along the screen;
[0074] In this embodiment of the invention, the ultrasonic emitting units to be activated can be further determined based on the calculated intersection positions. These ultrasonic emitting units are typically arranged along the edge of the screen or in a specific area to enhance the localization effect of the sound source and the auditory experience of the audience.
[0075] In some examples, the ultrasonic transmitting units closest to the intersection point can be calculated based on the location of the intersection point. These transmitting units are marked as target ultrasonic transmitting units and activated to emit ultrasonic signals of a specific frequency and intensity.
[0076] In some examples, the ultrasonic transmitting units can be symmetrically distributed on both sides of the perpendicular position of the intersection point. The specific positions can be selected as a square or circular area of a certain size above the intersection point.
[0077] Parametric array loudspeaker systems typically consist of multiple ultrasonic transmitting units, which are the key component responsible for emitting high-frequency ultrasonic signals. These units generally comprise ultrasonic transducers, drive circuits, control units, and a housing. These units are arranged in a specific geometric layout to form a parametric array. By controlling the emission signal of each ultrasonic transmitting unit and utilizing nonlinear acoustic effects, the high-frequency ultrasonic signal can be modulated into audible sound waves, thereby achieving functions such as directional audio transmission and noise control. By assembling ultrasonic transmitting units into a parametric array loudspeaker, high-precision, directional sound wave emission can be achieved, improving the performance and effectiveness in various application scenarios.
[0078] Step 104: Based on the location information of the audience, control the audio beams output by the multiple target ultrasonic transmitting units respectively.
[0079] In this embodiment of the invention, the audio beams output by multiple target ultrasonic transmitting units can be further controlled based on the audience's location information. For example, the direction of the audio beams can be controlled to face the audience, and the width of the audio beams can be controlled to cover the audience's head. This step aims to achieve a personalized audio experience, ensuring that each audience member receives the best auditory effect.
[0080] In some examples, the optimal auditory position for each viewer can be calculated based on their three-dimensional coordinates and movement trajectory. Then, the output parameters of the target ultrasonic transmitter, such as frequency, intensity, and phase, can be adjusted to ensure the audio beam is precisely directed to each viewer's location.
[0081] In this embodiment of the invention, the playback strategy can be dynamically adjusted according to the actual position of the audience and the position of the sound source image, thereby better adapting to different environments and user needs. By acquiring the sound source image position information and audience position information during the screen display process, real-time monitoring, precise positioning, and optimization of the interactive experience of the sound source and audience are achieved, thereby improving system efficiency and user experience, and helping users quickly solve the problem of audio-visual asynchrony.
[0082] Reference Figure 2 The diagram illustrates a flowchart of another sound source adjustment method based on audience position provided by an embodiment of the present invention. The method may specifically include the following steps:
[0083] Step 201: During the process of displaying the image on the screen, determine the location information of the sound source image;
[0084] In some embodiments, step 201 may include the following sub-steps:
[0085] Sub-step S21: When the sound source type is human, determine the center point position of the mouth of the human face in the sound image, and use the center point position of the mouth of the human face as the position information of the sound source image.
[0086] In this embodiment of the invention, when the sound source type is human, it indicates that the sound source image contains facial information. In this case, it is necessary to determine the center point of the mouth in the sound image and use this position as the location information of the sound source image.
[0087] In some examples, this method can be applied to intelligent audio-visual systems, which can use computer vision technology to identify and locate faces in audio-visual images. Specifically, the system uses a camera to capture audio-visual images and identifies facial regions in the images using face detection algorithms. Subsequently, the system further analyzes the facial regions to locate the center point of the mouth. This step typically involves facial landmark detection technology, which identifies key points such as the eyes, nose, and mouth to calculate the precise location of the mouth's center point.
[0088] In some examples, intelligent audio-visual systems can also incorporate deep learning models to improve the accuracy of face detection and mouth center point localization. For instance, intelligent audio-visual systems can utilize pre-trained face detection and keypoint detection models to process audio images, thereby achieving high-precision face and mouth center point localization. In this way, the system can accurately determine the location information of the sound source image, providing a basis for subsequent sound source localization and audience interaction.
[0089] Sub-step S22: When the sound source type is non-human, determine the horizontal center point position of the sound image and use the horizontal center point position as the position information of the sound source image.
[0090] In this embodiment of the invention, when the sound source type is non-human, it means that the sound source image does not contain facial information. In this case, it is necessary to determine the horizontal center point of the sound image and use this position as the location information of the sound source image.
[0091] In some examples, intelligent audio-visual systems can use image processing techniques to identify and locate the lateral center point of a sound image. Specifically, the system uses a camera to capture sound images and employs image processing algorithms to identify the sound source region within the image. Subsequently, the system further analyzes the sound source region to calculate the lateral center point of the image. This step typically involves image segmentation and edge detection techniques, identifying the boundaries of the sound source region to calculate the precise location of the lateral center point.
[0092] In some examples, the system can also incorporate machine learning models to improve the accuracy of sound source region identification and lateral center point localization. For instance, the system can utilize pre-trained image segmentation and edge detection models to process sound images, thereby achieving high-precision sound source region and lateral center point localization. In this way, the system can accurately determine the location information of the sound source image, providing a basis for subsequent sound source localization and audience interaction.
[0093] Step 202: Determine the location information of the audience;
[0094] In this embodiment of the invention, the intelligent audio-visual system can accurately identify and locate the audience's position information through various technical means, providing a basis for subsequent sound source localization and audience interaction.
[0095] In some examples, the system can capture audience location information in real time using multiple sensors such as cameras, ultrasonic sensors, and infrared sensors. Specifically, the system uses cameras to capture images of the audience and identifies their positions using computer vision technology. Simultaneously, the system can also use ultrasonic and infrared sensors to determine the audience's three-dimensional coordinates and movement trajectory by measuring the reflection time of sound waves and infrared radiation.
[0096] In some examples, the method for identifying audience positions can be as follows: cameras are distributed around or at the top of the screen to periodically capture images of the audience; the captured images are transmitted to the system's central processing unit (CPU) for preliminary processing and filtering; the system first determines a reference position, which is usually based on the center point of the screen or a preset comfortable viewing area; the system compares the position information of each audience member with the reference position to identify audience members whose positions are significantly deviated from or close to the reference position; if the position information of a certain audience member continues to deviate from the reference value, that is, is in the non-sweet spot area, the system will mark it as an abnormal audience position.
[0097] In some examples, cameras are positioned around the perimeter or top of the screen, periodically capturing images of viewers. This image data is transmitted to the system's central processing unit (CPU) for initial processing and filtering to remove noise and outliers. The system then determines a reference position based on the center point of the screen or a preset comfortable viewing area. Each viewer's position is compared to this reference position, identifying those whose positions significantly deviate from or are close to it. If a viewer's position consistently deviates from the reference value, the system marks them as an abnormal viewer position and triggers appropriate processing strategies.
[0098] By capturing and analyzing audience location information in real time using multiple sensors, the system can accurately identify audience positions and compare them with reference positions. It can identify audience positions that are significantly deviated from or close to the reference positions and mark them as abnormal audience positions, thereby triggering corresponding processing strategies.
[0099] Step 203: Determine the line connecting the location information of the sound source image and the location information of the audience, extending it to the intersection point on the screen;
[0100] In this embodiment of the invention, the intelligent audio-visual system accurately calculates the line connecting the sound source image and the audience position, and extends it to the intersection on the screen as a basis for subsequent judgment.
[0101] In some examples, the system can calculate the connection between a sound source image and an audience member by combining the latter's location information. Specifically, the system can use a microphone array to capture the location information of the sound source image and a camera or sensor network to capture the audience member's location information. The system can then match and calculate the location information of the sound source image and the audience member's location information, using geometric principles to determine the connection between the sound source and the audience member.
[0102] In some examples, the calculation method for the connection can be as follows: the location information of the sound source image is captured by a microphone array, and the location information of the audience is captured by a camera or sensor network; the captured data is transmitted to the system's central processing unit (CPU) for preliminary processing and filtering; the system first determines a reference position, which is usually based on the center point of the screen or a preset comfortable viewing area; the system matches and calculates the location information of the sound source image and the location information of the audience, and determines the connection between the sound source and the audience through geometric principles; the system extends the connection to the intersection on the screen and calculates the projection position.
[0103] In some instances, the location information of the sound source image is captured via a microphone array, while the audience's location information is captured via cameras or a sensor network. This data is transmitted to the system's central processing unit (CPU) for initial processing and filtering to remove noise and outliers. Subsequently, the system determines a reference position based on the center point of the screen or a preset comfortable viewing area. The system matches and calculates the location information of the sound source image with the audience's location information, using geometric principles to determine the line connecting the sound source and the audience. The system extends this line to its intersection on the screen to calculate the projection position.
[0104] By accurately calculating the line connecting the sound source image and the audience's position, and extending it to the intersection on the screen, the system can achieve precise sound source localization and audience interaction. By combining the positional information of the sound source image and the audience's positional information, calculating the line connecting the two, and extending it to the intersection on the screen, the system can identify the projection position of the sound source on the screen, thereby achieving a visual representation of the sound source.
[0105] Step 204: Determine the target area in the deployment area of the plurality of ultrasonic transmitting units along the vertical direction of the intersection point;
[0106] In this embodiment of the invention, the intelligent audio-visual system determines the target area in the deployment area of multiple ultrasonic transmitting units along the vertical direction of the intersection point.
[0107] In some examples, the system can determine the target area by combining the location information of the intersection points with the deployment area of the ultrasonic transmitting units. Specifically, the system uses the location information of the intersection points to calculate the range of the target area along the vertical direction. Subsequently, the system matches the deployment area of the ultrasonic transmitting units with the target area to identify the ultrasonic transmitting units located within the target area.
[0108] In some examples, the target area can be determined as follows: the location information of the intersection point is calculated by connecting the sound source image and the audience position; the system calculates the target area range along the vertical direction based on the location information of the intersection point; the deployment area of the ultrasonic transmitting unit is determined by a preset geometric layout; the system matches the deployment area of the ultrasonic transmitting unit with the target area to determine the ultrasonic transmitting unit located within the target area.
[0109] In some examples, the location information of the intersection point is calculated by connecting the sound source image and the audience's position. Based on the intersection point's location information, the system calculates the target area range along the vertical direction. The deployment area of the ultrasonic transmitting units is determined by a preset geometric layout, typically arranged along the screen edge or a specific area. The system matches the deployment area of the ultrasonic transmitting units with the target area to determine the ultrasonic transmitting units located within the target area.
[0110] In some examples, the ultrasonic transmitting units can be symmetrically deployed on both sides of the perpendicular line from the intersection point. The shape of the parametric array loudspeaker composed of ultrasonic transmitting units can be planned as a regular shape such as a square or circle of a predetermined size above the intersection point. For example, the parametric array loudspeaker can be a circle or square with a width equal to the width of a typical human head (18cm) or set to the observed head contour size of the audience. This patent does not limit the spacing between the ultrasonic transmitting units.
[0111] Step 205: The ultrasonic emitting unit located in the target area is identified as the target ultrasonic emitting unit;
[0112] In this embodiment of the invention, the intelligent audio-visual system identifies the ultrasonic transmitting unit located within the target area as the target ultrasonic transmitting unit, thereby achieving precise sound source localization and audience interaction.
[0113] In some examples, the system can identify target ultrasonic transmitters by combining the location information of the target area with the deployment area of the ultrasonic transmitters. Specifically, the system uses the location information of the target area to identify ultrasonic transmitters located within that area. Subsequently, the system marks these ultrasonic transmitters as target ultrasonic transmitters and performs appropriate control and management.
[0114] Step 206, based on the audience's location information, controlling the audio beams output by the plurality of target ultrasonic transmitting units respectively, including:
[0115] Sub-step S23: Based on the location information of the audience, control the direction of the audio beams output by the multiple target ultrasonic transmitting units to be towards the audience.
[0116] In this embodiment of the invention, the intelligent audio-visual system controls the direction of the audio beams output by multiple target ultrasonic transmitting units according to the audience's location information, so that the audio beams are directed toward the audience.
[0117] In some examples, the system can control the direction of the audio beam by combining the location information of the audience and the location information of the target ultrasonic transmitters. Specifically, the system uses the audience's location information to calculate the optimal direction of the audio beam output by each target ultrasonic transmitter. Subsequently, the system adjusts the output parameters of the target ultrasonic transmitters to ensure that the audio beam is precisely pointed at the audience's location.
[0118] By controlling the direction of the audio beams output by multiple target ultrasonic transmitting units based on the audience's location information, the system can achieve precise sound source localization and audience interaction. By combining the audience's location information and the location information of the target ultrasonic transmitting units to control the direction of the audio beams, the system can ensure that the audio beams are accurately pointed at the audience's location, thereby enhancing the audience's auditory experience.
[0119] Sub-step S24: Determine the distance of the plurality of target ultrasonic transmitting units relative to the audience based on the audience's location information; control the width of the audio beams output by the plurality of target ultrasonic transmitting units based on the distance of the plurality of target ultrasonic transmitting units relative to the audience.
[0120] In this embodiment of the invention, the intelligent audio-visual system determines the distance between multiple target ultrasonic transmitting units and the audience based on the audience's location information, and controls the width of the audio beam of each ultrasonic transmitting unit to ensure that the audio beam can cover the audience's head.
[0121] In some examples, the system can calculate the distance between each target ultrasonic transmitter and the audience by combining the audience's location information with the location information of the target ultrasonic transmitters. Specifically, the system uses the audience's location information to calculate the distance between each target ultrasonic transmitter and the audience. Then, based on these distances, the system adjusts the width of the audio beam output by the target ultrasonic transmitter to ensure that the audio beam covers the audience's location.
[0122] By determining the distances of multiple target ultrasonic transmitters relative to the audience based on their location information and controlling the width of the audio beam for each, the system achieves precise sound source localization and audience interaction. By combining the audience's location information with the location information of the target ultrasonic transmitters, the system calculates the distance of each transmitter relative to the audience and adjusts the audio beam width to ensure that the audio beam covers the audience's location, thereby enhancing their auditory experience.
[0123] Step 207: When the position information of the audience remains unchanged, but the position information of the sound source image changes, detect the change angle of the direction of the audio beams output by the multiple target ultrasonic transmitting units; when the change angle is greater than a preset angle threshold, redetermine the target ultrasonic transmitting units.
[0124] In this embodiment of the invention, the intelligent audio-visual system dynamically adjusts the target ultrasonic transmitting units by detecting the change angle of the direction of the audio beams output by multiple target ultrasonic transmitting units.
[0125] In some examples, the detection method for the change angle of the audio beam direction can be as follows: the audience's position information is captured by a camera or sensor network; the position information of the sound source image is captured by a microphone array; the system calculates the optimal direction of the audio beam output by each target ultrasonic transmitting unit based on the audience's position information; the system monitors the change in the position information of the sound source image and calculates the change angle of the audio beam direction; when the change angle is greater than a preset angle threshold, the system redetermines the target ultrasonic transmitting unit.
[0126] Reference Figure 3 As shown, the direction of the audio beam is defined as the azimuth angle θ of the viewer. The azimuth angle is the angle between the line connecting the viewer's position and the position of the sound source image and the perpendicular line to the screen. When the viewer's position remains unchanged but the position of the sound source image changes, a sound-image separation angle is generated. The sound-image separation angle is equal to the difference in azimuth angle before and after the change in the position of the sound source image, i.e., Δθ = θ1 - θ2. If the sound-image separation angle is too large, it will give the viewer a sense of sound-image separation. According to relevant literature, untrained viewers can perceive a sound-image deviation of more than 20°. Therefore, the preset angle threshold can be set to 20°. That is, when the viewer's azimuth angle changes by more than 20°, the system redetermines the intersection point based on the changed position of the sound source image, thereby determining the position of the target ultrasonic transmitting unit, the angle of the audio beam, and the coverage area.
[0127] Step 208: When the location information of the audience changes, but the location information of the sound source image remains unchanged, determine whether the location information of the audience exceeds the coverage range of the audio beams output by the multiple target ultrasonic transmitting units; when the location information of the audience exceeds the coverage range of the audio beams output by the multiple target ultrasonic transmitting units, then re-determine the target ultrasonic transmitting units.
[0128] In this embodiment of the invention, the intelligent audio-visual system detects changes in the audience's position information and determines whether the audience exceeds the coverage range of the audio beams output by multiple target ultrasonic transmitting units. If the audience exceeds the coverage range, the system dynamically adjusts the target ultrasonic transmitting units.
[0129] In some examples, the detection method for audience location information exceeding the coverage range of the audio beam can be as follows: the audience's location information is captured by a camera or sensor network; the location information of the target ultrasonic transmitting unit is determined by a preset geometric layout; the system calculates the coverage range of the audio beam output by each target ultrasonic transmitting unit based on the audience's location information and the distance between the audience and the target ultrasonic transmitting unit; the system monitors changes in the audience's location information and determines whether the audience exceeds the coverage range of the audio beam; when the audience's location information exceeds the coverage range of the audio beam, the system re-determines the target ultrasonic transmitting unit.
[0130] In some examples, the radiation angle α of the audio beam is related to factors such as the size of the parametric array speaker and the distance to the audience.
[0131] Reference Figure 4 As shown, for example, the constraint condition for the radiation angle α of the audio beam can be α≥2*arctan(d / 2D), where d is the width of the audience's head, which can be 18cm or the actual measured value, and D is the distance between the audience's head and the center of the parametric array speaker.
[0132] When the position of the sound source image remains unchanged but the audience position changes, the sound-image separation angle Δθ may be very small. However, if the audience leaves the audio beam coverage area of the original parametric array speaker, they will not hear the sound and will lose the sound-image synchronicity experience. In this case, the lateral displacement Δx of the audience in front of the screen and the coverage area of the audio beam are used as the basis for judgment. Since the radiation angle α of the audio beam is very small, the lateral width of the beam at the audience position is taken as tan(1 / 2α)*2D. When Δx > tan(1 / 2α)*D, that is, when the audience leaves the beam coverage area of the original parametric array speaker, the position of the parametric array speaker must be reselected and the new audience azimuth angle θ2 is used as the beam direction to emit sound.
[0133] In this embodiment of the invention, the playback strategy can be dynamically adjusted according to the actual position of the audience and the position of the sound source image, thereby better adapting to different environments and user needs. By acquiring the sound source image position information and audience position information during the screen display process, and setting a preset threshold to dynamically adjust the target ultrasonic transmitting unit, real-time monitoring, precise positioning, and optimization of the interactive experience of the sound source and audience are achieved, thereby improving system efficiency and user experience, and helping users quickly solve the problem of audio-visual asynchrony.
[0134] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0135] Reference Figure 5 The diagram illustrates a structural block diagram of a sound source adjustment device based on audience position according to an embodiment of the present invention, which may specifically include the following modules:
[0136] The position information determination module 301 is used to determine the position information of the sound source image and the position information of the audience during the process of displaying the image on the screen.
[0137] The intersection point determination module 302 is used to determine the intersection point of the line connecting the position information of the sound source image and the position information of the audience, extending to the screen.
[0138] The target ultrasonic transmitting unit determination module 303 is used to determine multiple target ultrasonic transmitting units from multiple ultrasonic transmitting units arranged along the screen based on the intersection point.
[0139] The audio beam control module 304 is used to control the audio beams output by the multiple target ultrasonic transmitting units according to the position information of the audience.
[0140] In this embodiment of the invention, the intersection point determination module includes:
[0141] The sound source type determination submodule is used to determine the sound source type;
[0142] The first location information determination submodule is used to determine the center point position of the mouth of a human face in the sound image when the sound source type is human, and to use the center point position of the mouth of the human face as the location information of the sound source image.
[0143] The second location information determination submodule is used to determine the horizontal center point position of the sound image when the sound source type is non-human, and use the horizontal center point position as the location information of the sound source image.
[0144] Optionally, the plurality of ultrasonic emitting units are arranged along the lateral edge of the screen; the target ultrasonic emitting unit determining module includes:
[0145] The target area determination submodule is used to determine the target area along the vertical direction of the intersection point within the deployment area of the plurality of ultrasonic transmitting units;
[0146] The target ultrasonic transmitting unit submodule is used to identify ultrasonic transmitting units located within the target area as target ultrasonic transmitting units.
[0147] In this embodiment of the invention, the audio beam control module includes:
[0148] The audio beam direction control submodule is used to control the direction of the audio beams output by the plurality of target ultrasonic transmitting units toward the audience, based on the audience's position information.
[0149] In this embodiment of the invention, the audio beam control module further includes:
[0150] The distance determination submodule is used to determine the distance between the plurality of target ultrasonic transmitting units and the audience based on the audience's location information;
[0151] The width determination submodule is used to control the width of the audio beams output by the multiple target ultrasonic transmitting units based on the distances of the multiple target ultrasonic transmitting units relative to the audience.
[0152] In this embodiment of the invention, the method further includes:
[0153] An angle detection module is used to detect the change angle of the direction of the audio beams output by the multiple target ultrasonic transmitting units when the position information of the audience remains unchanged and the position information of the sound source image changes.
[0154] The first re-determination submodule is used to re-determine the target ultrasonic transmitting unit when the changed angle is greater than a preset angle threshold.
[0155] In this embodiment of the invention, the method further includes:
[0156] The coverage area determination module is used to determine whether the location information of the audience exceeds the coverage area of the audio beams output by the multiple target ultrasonic transmitting units when the location information of the audience changes but the location information of the sound source image remains unchanged.
[0157] The second re-determination submodule is used to re-determine the target ultrasonic transmitting unit when the location information of the audience exceeds the coverage range of the audio beams output by the plurality of target ultrasonic transmitting units.
[0158] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0159] This invention also provides an electronic device, comprising:
[0160] It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When executed by the processor, the computer program implements the various processes of the above-described embodiment of the sound source adjustment method based on audience position and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0161] This invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described sound source adjustment method embodiment based on audience position and achieves the same technical effect. To avoid repetition, it will not be described again here.
[0162] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0163] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0164] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0165] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0166] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0167] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0168] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0169] The present invention has provided a detailed description of a sound source adjustment method, apparatus, device, and medium based on audience position. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A sound source adjustment method based on audience position, characterized in that, include: During the process of displaying images on the screen, the location information of the sound source image and the location information of the audience are determined; Determine the location information of the sound source image and the location information of the audience, and extend the line connecting them to the intersection on the screen; Based on the intersection point, a plurality of target ultrasonic transmitting units are determined from a plurality of ultrasonic transmitting units arranged along the screen; Based on the location information of the audience, the audio beams output by the multiple target ultrasonic transmitting units are controlled respectively; When the position information of the audience remains unchanged, but the position information of the sound source image changes, the angle of change in the direction of the audio beams output by the multiple target ultrasonic transmitting units is detected. When the change angle is greater than a preset angle threshold, the target ultrasonic transmitting unit is redefined.
2. The method according to claim 1, characterized in that, The determination of the location information of the sound source image includes: Determine the type of sound source; When the sound source type is human, determine the position of the center point of the mouth of the human face in the sound source image, and use the position of the center point of the mouth of the human face as the position information of the sound source image; When the sound source type is non-human, the horizontal center point position of the sound source image is determined, and the horizontal center point position is used as the position information of the sound source image.
3. The method according to claim 1, characterized in that, The plurality of ultrasonic emitting units are arranged along the lateral edge of the screen; the step of determining a plurality of target ultrasonic emitting units from the plurality of ultrasonic emitting units arranged along the screen based on the intersection point includes: Along the vertical direction of the intersection point, a target area is determined within the deployment area of the plurality of ultrasonic transmitting units; The ultrasonic transmitting unit located within the target area is identified as the target ultrasonic transmitting unit.
4. The method according to claim 1, characterized in that, The step of controlling the audio beams output by the plurality of target ultrasonic transmitting units according to the audience's location information includes: Based on the location information of the audience, the direction of the audio beams output by the plurality of target ultrasonic transmitting units is controlled to be towards the audience.
5. The method according to claim 4, characterized in that, The step of controlling the audio beams output by the plurality of target ultrasonic transmitting units according to the audience's location information further includes: Based on the location information of the audience, the distances of the plurality of target ultrasonic transmitting units relative to the audience are determined; The width of the audio beams output by the multiple target ultrasonic transmitting units is controlled according to the distance between the multiple target ultrasonic transmitting units and the audience.
6. The method according to claim 1, characterized in that, Also includes: When the location information of the audience changes, but the location information of the sound source image remains unchanged, determine whether the location information of the audience exceeds the coverage range of the audio beams output by the multiple target ultrasonic transmitting units; When the location information of the audience exceeds the coverage range of the audio beams output by the plurality of target ultrasonic transmitting units, the target ultrasonic transmitting units are re-determined.
7. A sound source adjustment device based on audience position, characterized in that, The device includes: The location information determination module is used to determine the location information of the sound source image and the location information of the audience during the process of displaying the image on the screen. The intersection point determination module is used to determine the intersection point of the line connecting the position information of the sound source image and the position information of the audience, extending to the screen. A target ultrasonic transmitting unit determination module is used to determine multiple target ultrasonic transmitting units from multiple ultrasonic transmitting units arranged along the screen based on the intersection point; An audio beam control module is used to control the audio beams output by the multiple target ultrasonic transmitting units according to the position information of the audience. An angle detection module is used to detect the change angle of the direction of the audio beams output by the multiple target ultrasonic transmitting units when the position information of the audience remains unchanged and the position information of the sound source image changes. The first re-determination submodule is used to re-determine the target ultrasonic transmitting unit when the changed angle is greater than a preset angle threshold.
8. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the steps of the sound source adjustment method based on audience position as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the sound source adjustment method based on audience position as described in any one of claims 1-6.