Directional sound transmission method and device based on visual aiming

CN122598601APending Publication Date: 2026-08-18POWER IDEA TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610757387.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-28
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]依赖预注册信息:基于人脸识别或身份识别的方法需要预先建立数据库,无法实现“即指即用”,且在公共场所存在隐私合规风险

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122598601A_ABST
    Figure CN122598601A_ABST
Patent Text Reader

Abstract

The application discloses a directional sound transmission method based on visual aiming, which comprises the following steps: acquiring a real-time scene image in an aiming direction; identifying a target object in the real-time scene image located in a preset aiming mark area, and generating spatial position information of the target object, wherein the preset aiming mark is a fixed pattern superimposed on a viewfinder interface; controlling an ultrasonic phased array transducer array to generate a parametric array sound beam propagating along a specified direction according to the spatial position information; and modulating an audio signal to be transmitted to an ultrasonic carrier of the parametric array sound beam, so that the audio signal is audible only at the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of directional audio transmission technology, and more specifically, to a directional sound transmission method and apparatus based on visual aiming. Background Technology

[0002] The need for private one-on-one or one-to-many communication in public places, while avoiding overhearing by unauthorized individuals, has been a long-standing technological requirement. Traditional solutions include the use of headsets, walkie-talkies, and other similar devices. However, these solutions require both parties to wear dedicated terminals, which presents inconveniences in temporary or non-fixed-team scenarios (such as museum tours, guided tours, and temporary group activities).

[0003] Existing directional sound transmission technologies mainly fall into three categories. The first category uses UWB tag positioning: the target needs to wear a tag, and the system tracks and drives the pan-tilt unit to align via the tag signal. The limitation of this approach is that it requires the target to wear a UWB tag, thus restricting deployment in temporary or random communication scenarios. The second category uses light source marker positioning: the target is equipped with an active light source, and the camera guides the sound beam after detecting the light source's position. This approach also requires the target to carry auxiliary equipment, and positioning accuracy may be affected under strong ambient light or light source obstruction conditions. The third category is directional content delivery technology based on facial analysis and target attribute classification: the system identifies the target's facial features and matches them with preset attributes or categories, then broadcasts corresponding content to the target accordingly. This type of technology usually requires pre-established databases or classification rules, making it difficult to handle targets that are not pre-registered, and its core lies in content matching based on attribute tags, rather than instant target locking.

[0004] It is evident that while existing directional sound transmission technologies have made significant progress in controlling the directionality of sound beams, they still have the following shortcomings in the interactive aspect of "how to conveniently and naturally select communication targets": Reliance on auxiliary equipment: It requires the target to wear tags, light sources and other auxiliary markers, which increases the cost of use and the difficulty of deployment, and is not suitable for random and temporary communication scenarios.

[0005] Reliance on pre-registration information: Methods based on facial recognition or identity verification require the pre-establishment of a database, which cannot achieve "point-and-use" functionality and poses privacy compliance risks in public places.

[0006] The interaction is not intuitive enough: the existing solutions have a high operating threshold and lack an intuitive interaction method that allows you to select a target by "pointing with your hand" or "looking with your eyes".

[0007] Therefore, how to provide a method for selecting directional sound transmission targets that requires no auxiliary marking or pre-registration and is intuitive to operate, and how to achieve accurate directional audio transmission on this basis, is a technical problem that urgently needs to be solved in this field.

[0008] It is important to note that the techniques described in this section are not necessarily those previously conceived or adopted. Unless otherwise specified, no technique described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be recognized in any prior art. Summary of the Invention

[0009] This application provides a directional sound transmission method and device based on visual aiming, which adopts "what you see is what you select" visual aiming interaction to accurately directionally transmit sound. It has the advantages of low threshold for use, flexible pointing and good privacy.

[0010] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A directional sound transmission method based on visual aiming includes the following steps: S100, acquiring a real-time scene image in the aiming direction; S200, identifying a target object located within a preset aiming mark area in the real-time scene image and generating spatial position information of the target object, wherein the preset aiming mark is a fixed graphic superimposed on the viewfinder; S300, controlling an ultrasonic phased array transducer array to generate a parametric array sound beam propagating along a specified direction according to the spatial position information; S400, modulating the audio signal to be transmitted onto the ultrasonic carrier of the parametric array sound beam so that the audio signal is audible only at the target object.

[0011] Preferably, step S200 further includes: when there are multiple candidate objects within the preset aiming mark area, calculating the depth distance and the area of ​​the face facing the camera of each candidate object in the image frame, and taking the one with the closest depth distance and the largest area of ​​the face facing the camera as the target object.

[0012] Preferably, step S200 further includes: if no facial information is detected, then a complete human body outline is detected as the target object.

[0013] Preferably, in step S200, the target object located within the preset aiming mark area in the real-time scene image is identified based on the convolutional neural network, and the target pixel coordinates in the two-dimensional image of the real-time scene image are converted into three-dimensional spatial coordinates by combining monocular depth estimation or binocular vision ranging algorithm to obtain the spatial position information of the target object.

[0014] Preferably, the spatial location information includes the azimuth, pitch, and distance of the target object relative to the detection location.

[0015] Preferably, step S300 is: controlling the driving phase difference of each transducer unit in the ultrasonic phased array transducer array according to the spatial position information to generate the parametric array acoustic beam in the specified direction.

[0016] Preferably, the method further includes the step: S500, when the target object moves, the spatial position information of the target object is updated in real time, and the ultrasonic phased array transducer array is controlled to synchronously adjust the beam pointing angle to keep the beam always aligned with the target object.

[0017] Preferably, the method further includes the step: S600, identifying whether there are non-target personnel within a preset range around the target object based on the real-time scene image; if there are non-target personnel, reducing the transmission power of the parametric array sound beam to below a preset safety threshold, or interrupting the transmission of the parametric array sound beam.

[0018] This invention also discloses a directional sound transmission device based on visual aiming. The device includes an image acquisition module, an aiming display module, a target recognition and positioning module, an ultrasonic phased array transducer array, and an audio input and modulation module. The image acquisition module is used to acquire real-time scene images in the aiming direction. The aiming display module is used to display the real-time scene image and preset aiming marks superimposed on the real-time scene image. The target recognition and positioning module is used to identify target objects located within the preset aiming mark area in the real-time scene image and generate spatial position information of the target objects. The ultrasonic phased array transducer array is used to generate a parametric array sound beam propagating in a specified direction based on the spatial position information. The audio input and modulation module is used to acquire the audio signal to be transmitted and modulate the audio signal onto the ultrasonic carrier of the parametric array sound beam so that the audio signal is audible only at the target object.

[0019] Preferably, the device further includes a target confirmation button and a launch status indicator module. The target confirmation button is used to activate the target recognition and positioning module to lock the target and highlight it on the display screen when the user half-presses the button, and to start transmitting the modulated audio beam when the button is fully pressed. The launch status indicator module is used to generate a visual prompt on the aiming display module simultaneously when the ultrasonic phased array transducer array transmits the parametric array sound beam.

[0020] It should be understood that the description in this section is not intended to identify key or important features of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0021] The accompanying drawings exemplify embodiments and form part of the specification, working together with the textual description to explain exemplary implementations of the embodiments. The drawings shown are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0022] Figure 1 A flowchart of a vision-targeting-based directional sound transmission method according to a preferred embodiment of the present invention; Figure 2 This is a schematic diagram of the display interface of a target object within a preset aiming mark area according to a preferred embodiment of the present invention; Figure 3 This is a schematic diagram of the acoustic beam pointing control principle of an ultrasonic phased array transducer array according to a preferred embodiment of the present invention. Figure 4 A flowchart of a vision-targeting-based directional sound transmission method according to another preferred embodiment of the present invention; Figure 5 This is a flowchart of a vision-targeting-based directional sound transmission method according to another preferred embodiment of the present invention; Figure 6 A block diagram of a vision-aiming directional sound transmission device according to a preferred embodiment of the present invention; Figure 7 This is a schematic diagram of the usage status of a museum interpretation scenario provided in Embodiment 1 of the present invention; Figure 8 This is a schematic diagram illustrating the usage status of a private communication scenario in a noisy environment, as provided in Embodiment 2 of the present invention. Figure 9 This is a structural schematic diagram of the smart glasses form provided in Embodiment 3 of the present invention. Detailed Implementation

[0023] To make the inventive objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0025] In the description of the embodiments of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. The term "multiple" means two or more, unless otherwise explicitly specified. The term "comprising" indicates the presence of the described feature, whole, step, operation, element, and / or component, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or sets thereof. The term "and / or" describes the relationship between related objects, indicating that three relationships may exist. For example, A and / or B may include three cases: A existing alone, A and B existing simultaneously, and B existing alone. The character " / " generally indicates that the related objects before and after are in an "or" relationship.

[0026] Unless otherwise defined, all technical terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art; the terms used in the embodiments of this application are for the purpose of describing specific embodiments only and are not intended to limit this application; the terms "comprising" and "having" and any variations thereof in the specification, claims and foregoing description of the drawings of this application are intended to cover non-exclusive inclusion.

[0027] Furthermore, terms such as "exemplary," "for example," and "optional" are used to indicate illustrative purposes. Any technical solution described by the above terms in the embodiments of this application should not be construed as being more preferred or advantageous than other technical solutions. Specifically, these terms are intended to present the relevant technical concepts in terms of specific implementation methods.

[0028] Figure 1The flowchart of a preferred embodiment of a vision-based directional sound transmission method according to the present invention includes the following steps: S100, acquiring a real-time scene image in the aiming direction; S200, identifying a target object located within a preset aiming mark area in the real-time scene image and generating spatial position information of the target object, wherein the preset aiming mark is a fixed graphic superimposed on the viewfinder; S300, controlling an ultrasonic phased array transducer array to generate a parametric array sound beam propagating along a specified direction based on the spatial position information; and S400, modulating the audio signal to be transmitted onto the ultrasonic carrier of the parametric array sound beam to ensure that the audio signal is audible only at the target object. It is understood that the specified direction refers to the direction from the ultrasonic phased array transducer array to the target object. Specifically, the viewfinder typically refers to the display interface of the scene image.

[0029] This invention eliminates reliance on auxiliary equipment by employing a purely visual aiming and locking mechanism. Target objects do not require any tags, light sources, or other auxiliary markers, significantly lowering the deployment threshold for temporary communication scenarios. It allows for immediate use without pre-registration; target recognition does not rely on any pre-established facial or identity databases. Users simply align the aiming marker on the viewfinder with the target, effectively protecting the privacy of non-target individuals in public places. The interaction is intuitive and natural; the "what you see is what you select" aiming method makes user operation as natural as taking a photo or shining a flashlight, with virtually zero learning cost. This invention utilizes "what you see is what you select" visual aiming interaction for precise directional sound transmission, offering advantages such as low barrier to entry, flexible pointing, and good privacy. It can be widely applied in museum explanations, private communication in noisy environments, hearing aids, and military command.

[0030] In specific implementations, the preset aiming mark can be a fixed graphic such as a crosshair, circle, or frame. The aiming direction is moved so that the target object to be communicated with appears within the area of ​​the preset aiming mark, and faces, body outlines, or other preset types of target objects within this area are identified. Figure 2 The figure shows a schematic diagram of the display interface of the target object located in the preset aiming mark area in the real-time scene image identified in the technical solution of the present invention. The aiming mark in the figure is a crosshair or a circular frame. The locked target object is outlined with a highlighted border, showing the outline of a face or human body. At the same time, the status prompt text "Target locked" is displayed, and status icons such as battery / volume are displayed in the lower right corner.

[0031] In a preferred embodiment, step S200 may further include: when multiple candidate objects exist within the preset aiming marker area, calculating the depth distance and the area of ​​each candidate object facing the camera in the image frame, and selecting the one with the closest depth distance and the largest area of ​​the face facing the camera as the target object. Multiple candidate objects refer to multiple detected objects appearing in the viewfinder, such as multiple faces or multiple human silhouettes. Specifically, if no facial information is detected, then detecting a complete human silhouette as the target object. This strategy avoids the loss of target objects due to facial occlusion, such as wearing a mask or turning the head, and improves the locking robustness in complex scenes.

[0032] In a preferred embodiment, step S200 can identify target objects located within a preset aiming mark area in a real-time scene image based on a convolutional neural network (e.g., YOLO series, SSD, etc.), and combine this with monocular depth estimation or binocular vision ranging algorithms to convert the target pixel coordinates in the two-dimensional image of the real-time scene image into three-dimensional spatial coordinates, thereby obtaining the spatial position information of the target object. In a specific embodiment, the spatial position information of the target object, i.e., the aforementioned three-dimensional spatial coordinates, may include the azimuth angle, pitch angle, and distance of the target object relative to the detection position.

[0033] In a preferred embodiment, step S300 can be: controlling the driving phase difference of each transducer unit in the ultrasonic phased array transducer array according to the spatial position information to generate a parametric array sound beam in the specified direction. This enables electronic pointing control of the sound beam without mechanical rotation, allowing the sound beam turning speed to reach millisecond levels, and enabling it to follow rapid changes in the user's aiming direction in real time. Figure 3 This diagram illustrates the principle of beam pointing control for an ultrasonic phased array transducer array. The controller (DSP) controls phase shifters 1 to 15, with a phase difference... Δ φ is successively from 0 to (15-1)δ, resulting in phases 1 to 16, which control phase shifters 1 to 16 respectively, thereby obtaining a parametric array acoustic beam with a pointing angle of θ.

[0034] In a preferred embodiment, after necessary preprocessing, such as compression and equalization, the audio signal can be loaded onto the ultrasonic carrier of the parametric array sound beam using amplitude modulation or single-sideband modulation. When the modulated ultrasonic beam propagates through the air, it self-demodulates through the nonlinear acoustic effects of air to produce a highly directional audible sound. The energy of this audible sound is mainly concentrated along the propagation path of the sound beam, and attenuates rapidly outside the path. Therefore, only the targeted object can hear it clearly, and bystanders, even at very close distances, will find it difficult to detect.

[0035] In a preferred embodiment, such as Figure 4As shown, the directional sound transmission method based on visual aiming of the present invention may further include step S500: when the target object moves, the spatial position information of the target object is updated in real time, and the ultrasonic phased array transducer array is controlled to synchronously adjust the beam pointing angle to keep the beam always aligned with the target object. That is, after locking onto the target object, the positional changes of the target object in consecutive image frames can be continuously tracked. Since the electronic adjustment speed of the phased array beam is much faster than that of a mechanical pan-tilt head, the present invention can effectively handle scenarios where the target moves laterally at a speed of up to 1.5 m / s within a distance of 2 m, and is suitable for communication between people moving around in exhibitions. It can continuously follow the moving target and synchronously adjust the beam direction, ensuring communication continuity in moving scenarios and enhancing dynamic tracking continuity.

[0036] In a preferred embodiment, such as Figure 5 As shown, the directional sound transmission method based on visual aiming of the present invention may further include step S600: identifying whether there are non-target personnel within a preset range around the target object based on the real-time scene image; if non-target personnel are present, reducing the transmission power of the parametric array sound beam to below a preset safety threshold, or interrupting the transmission of the parametric array sound beam. A warning may also be issued simultaneously, and monitoring may continue until the danger is eliminated. Automatic power reduction or interruption when non-target personnel are detected approaching the sound beam path avoids causing discomfort or hearing damage to unrelated personnel. Specifically, if no non-target personnel are present, normal power transmission is maintained. This invention also discloses a directional sound transmission device based on visual aiming, such as... Figure 6 As shown, the device includes an image acquisition module 100, an aiming display module 200, a target recognition and positioning module 300, an ultrasonic phased array transducer array 400, and an audio input and modulation module 500. The image acquisition module 100 is used to acquire real-time scene images in the aiming direction; the aiming display module 200 is used to display the real-time scene image and preset aiming marks superimposed on the real-time scene image; the target recognition and positioning module 300 is used to identify target objects located within the preset aiming mark area in the real-time scene image and generate spatial position information of the target objects; the ultrasonic phased array transducer array 400 is used to generate a parametric array sound beam propagating in a specified direction based on the spatial position information; the audio input and modulation module 500 is used to acquire the audio signal to be transmitted and modulate the audio signal onto the ultrasonic carrier of the parametric array sound beam so that the audio signal is audible only at the target object.

[0037] The directional sound transmission device based on vision aiming of the present invention can be a dedicated directional sound transmission device, or a smartphone, tablet computer, or smart glasses integrated with a camera. The image acquisition module 100 includes at least one camera for acquiring real-time scene images in the aiming direction. Preferably, it can include a color camera and a depth camera, such as a ToF camera or a binocular camera, to improve spatial positioning accuracy. Specifically, the real-time scene image in the aiming direction can be displayed through the device's viewfinder, which can be the device's display screen or the field of view of a head-mounted display. When the user points the device in a certain direction, the real-time scene image in that direction is displayed in the viewfinder.

[0038] The directional sound transmission device based on vision aiming of the present invention typically also includes a housing, and the image acquisition module 100 can be disposed at the center of the front end of the housing. The aiming display module 200 can be an LCD screen embedded in the rear end of the housing, facing the user, or it can be an external display wirelessly connected to the device. The target recognition and positioning module 300 can be a program module run by a processor, and can include a face / human detection submodule and a spatial mapping submodule. The face / human detection submodule can recognize faces or complete human contours in images based on convolutional neural networks; the spatial mapping submodule can combine monocular depth estimation or binocular visual ranging algorithms to convert the target pixel coordinates in a two-dimensional image into three-dimensional spatial coordinates (azimuth, pitch, distance), thereby obtaining the spatial position information of the target object. The ultrasonic phased array transducer array 400 can be disposed at the front end or top of the housing, and consists of multiple ultrasonic transducer units, which can be arranged in a ring, rectangular or spherical array for generating and emitting directional parametric array sound beams. The audio input and modulation module 500 can include a microphone or audio input interface, and can be disposed on the side or top of the housing.

[0039] In a preferred embodiment, the directional sound transmission device based on visual aiming of the present invention may further include a target confirmation button and a transmission status indicator module. The target confirmation button is used to activate the target recognition and positioning module to lock onto the target and highlight it on the display screen when the user half-presses the button; when the button is fully pressed, the modulated audio beam is emitted. The transmission status indicator module is used to simultaneously generate a visual prompt on the aiming display module when the ultrasonic phased array transducer array emits the parametric array sound beam. For example, displaying a red halo or flashing icon at the edge of the screen can alert onlookers that a directional sound transmission operation is currently underway, enhancing the social acceptability of the technology. The target confirmation button can be located in the grip area of ​​the housing, within thumb reach, for convenient user operation.

[0040] The following examples illustrate the implementation of the intelligent guide system for museums / art galleries, private communication in noisy environments, and smart glasses using the technical solution of this invention: Example 1: Smart Guidebook for Museums / Art Galleries This embodiment demonstrates the application of the present invention in intelligent interpretation scenarios in museums or art galleries.

[0041] Visitor A enters the exhibition hall and receives a directional audio guide device based on this invention. The device (handheld device) resembles a telescope without a lens, with a built-in camera and ultrasonic phased array transducer array at the front, and a handle and display screen at the rear.

[0042] (1) Visitors hold the audio guide to their eyes and observe the exhibits (famous paintings) in the exhibition hall through the display screen.

[0043] (2) A circular aiming mark is superimposed in the center of the display screen. Visitors move the audio guide and aim the mark at a painting.

[0044] (3) The built-in target recognition module of the guide identifies the target in the aiming area as the famous painting "Exhibit A" in about 0.2 seconds and selects it with a highlighted border on the screen.

[0045] (4) When the visitor presses the confirmation button on the handle, the audio guide will automatically play an audio explanation about exhibit A. The audio signal is modulated onto an ultrasonic carrier wave and directionally transmitted to exhibit A through a parametric array sound beam (dashed cone or beam).

[0046] The sound beam coverage was precisely controlled within an area of ​​approximately 30cm radius centered on the exhibit. For example... Figure 7 As shown, visitor B, standing next to the famous painting exhibit A, can clearly hear the explanation, while visitor C, standing on the other side of the same display case, can hear almost nothing. This effect stems from the strong directivity of the parametric array sound beam—the sidelobe energy is suppressed to below -20dB, so even if two people stand side by side, the one who is not aligned with the beam cannot obtain comprehensible information.

[0047] When visitor A moves to the next exhibit B, the above steps are repeated without any additional operation.

[0048] The smart audio guide for museums / art galleries in this embodiment has a working distance of 0.5m - 5m; a sound beamwidth of -6dB, approximately a circular area with a diameter of 15cm at a distance of 1m; and a maximum output sound pressure level of 85dB, adjustable, meeting hearing safety standards. Compared to traditional wireless headphones, this audio guide eliminates the process of borrowing and returning headphones; compared to public speakers, it avoids mutual interference caused by multiple exhibits sounding simultaneously in the exhibition hall; and compared to scanning QR codes for audio explanations, it eliminates the cumbersome steps of taking out a phone, scanning a code, and selecting playback. This achieves an immersive visiting experience of "one explanation per exhibit, without mutual interference."

[0049] Example 2: Private Communication in Noisy Environments (Street / Exhibition) This example demonstrates an application for private communication between two people in a noisy environment.

[0050] User A and User B are at an exhibition where the noise level is as high as 85dB. A wants to say something to B that only B can hear, such as, "Let's go to the next booth, it's too crowded here."

[0051] (1) User A takes out a small directional microphone (similar in shape to a wireless microphone receiver, with a camera and ultrasonic array at the front end) from his pocket.

[0052] (2) User A points the device toward B and observes through the small screen on the device. A crosshair mark is displayed in the center of the screen.

[0053] (3) User A moves a mobile device and aims the crosshair at B's face. The device automatically detects B's facial features, without needing to identify him, but simply determines that he is a valid target, and highlights B's head area on the screen, while vibrating to indicate "Target locked".

[0054] (4) User A speaks into the built-in microphone of the device, and the sound is modulated into a parametric array sound beam in real time and emitted only in the direction of B.

[0055] (5) B can clearly hear A's voice from about 2 meters away, as if A is whispering in his ear. User C, who is standing next to A, is not locked and cannot hear any sound at all. This is because the sidelobe sound pressure level of the parametric array beam at a distance of 1 meter is lower than the ambient noise floor (measured to be lower than 45dB), which is much lower than the threshold of human hearing.

[0056] (6) When user B moves, the device continuously locks onto B and the direction of the sound beam automatically adjusts accordingly, maintaining the continuity of communication.

[0057] The effective communication range of this small directional microphone is 2m-8m; the signal-to-noise ratio received by B is not less than 15dB in an ambient noise environment of 85dB; and the battery lasts for 4 hours of continuous use. Compared to mobile phone calls, there is no need to dial or for the other party to answer; compared to speaking close to the ear, it maintains social distance, is more hygienic, and less conspicuous; and compared to walkie-talkies, the content cannot be eavesdropped on by others. This method is particularly suitable for exhibition guidance, internal security communication, and assistive dialogue for hearing-impaired individuals in extremely noisy environments.

[0058] Example 3: Smart Glasses Form (Hands-free) This embodiment demonstrates a wearable application that integrates the present invention into augmented reality glasses.

[0059] The user wears a pair of AR smart glasses, which integrate a front-facing camera that can be set in the center of the frame or the front of the temples, an AR waveguide display located in the lens area, a battery / chip that can be set inside the temples, and an ultrasonic phased array transducer array integrated on the top of the frame that can be set on the top of the frame or on both sides of the temples. The glasses also include temples, which can have built-in circuitry.

[0060] (1) Users observe the world through the natural field of vision of AR glasses, and a semi-transparent virtual crosshair mark is always displayed in the center of the field of vision.

[0061] (2) The user turns his head and aims the crosshair at the target, which can be a person, a smart speaker, a service robot or a public information screen.

[0062] (3) When the crosshair is aligned with the target and stays for 0.5 seconds, the glasses will automatically lock onto the target and prompt "[target name] locked" through bone conduction headphones or directional sound beam.

[0063] (4) When the user issues a voice command, such as “broadcast today’s news”, or presses the shortcut key on the temple, the built-in audio module of the glasses modulates the content into a directional sound beam and transmits it only to the locked target.

[0064] The differentiating features of this smart glasses compared to the above embodiments are: Speaking to yourself: After the user locks the public information screen, the glasses transmit audio directionally to the user's own ear, enabling private information access that is "heard wherever you look," and cannot be heard by others.

[0065] Speak to others: After a user locks onto another user wearing the same type of glasses, they can have a private conversation directly through the glasses without having to operate their phone.

[0066] This embodiment upgrades the target selection action from "handheld focusing" to "natural gaze," adding almost no cognitive burden. It frees up the hands, making it suitable for scenarios requiring continuous two-handed operation or where holding the device is inconvenient, such as outdoor sports, industrial maintenance, and military operations. Simultaneously, because the sound beam is also directional, even if someone is standing nearby, they cannot eavesdrop, ensuring information security and social etiquette.

[0067] The above embodiments are only used to illustrate the present application and are not intended to limit it. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included within the protection scope of the present application.

Claims

1. A directional sound transmission method based on visual aiming, characterized in that, Including the following steps: S100, acquire real-time scene images in the aiming direction; S200, identify the target object located within the preset aiming mark area in the real-time scene image, and generate the spatial position information of the target object, wherein the preset aiming mark is a fixed graphic superimposed on the viewfinder; S300, based on the spatial position information, control the ultrasonic phased array transducer array to generate a parametric array sound beam that propagates along a specified direction; S400, the audio signal to be transmitted is modulated onto the ultrasonic carrier of the parametric array sound beam so that the audio signal can be heard only at the target object.

2. The directional sound transmission method based on visual aiming according to claim 1, characterized in that, Step S200 further includes: When there are multiple candidate objects within the preset aiming mark area, the depth distance and the area of ​​the face facing the camera of each candidate object in the image frame are calculated, and the one with the closest depth distance and the largest area of ​​the face facing the camera is taken as the target object.

3. The directional sound transmission method based on visual aiming according to claim 2, characterized in that, Step S200 further includes: If no facial information is detected, the complete human body outline is detected as the target object.

4. The directional sound transmission method based on visual aiming according to claim 1, characterized in that, Step S200 identifies the target object located within the preset aiming mark area in the real-time scene image based on a convolutional neural network, and combines monocular depth estimation or binocular vision ranging algorithm to convert the target pixel coordinates in the two-dimensional image of the real-time scene image into three-dimensional spatial coordinates to obtain the spatial position information of the target object.

5. The directional sound transmission method based on visual aiming according to claim 4, characterized in that, The spatial location information includes the azimuth, pitch, and distance of the target object relative to the detection location.

6. The directional sound transmission method based on visual aiming according to claim 1, characterized in that, Step S300 is as follows: control the driving phase difference of each transducer unit in the ultrasonic phased array transducer array according to the spatial position information to generate the parametric array acoustic beam in the specified direction.

7. The directional sound transmission method based on visual aiming according to any one of claims 1-6, characterized in that, The method further includes the following steps: S500: When the target object moves, the spatial position information of the target object is updated in real time, and the ultrasonic phased array transducer array is controlled to synchronously adjust the beam pointing angle to keep the beam always aligned with the target object.

8. The directional sound transmission method based on visual aiming according to any one of claims 1-6, characterized in that, The method further includes the following steps: S600: Based on the real-time scene image, identify whether there are non-target personnel within a preset range around the target object. If there are non-target personnel, reduce the transmission power of the parametric array sound beam to below a preset safety threshold, or interrupt the transmission of the parametric array sound beam.

9. A directional sound transmission device based on vision aiming, characterized in that, The device includes an image acquisition module, an aiming and display module, a target recognition and positioning module, an ultrasonic phased array transducer array, and an audio input and modulation module. The image acquisition module is used to acquire real-time scene images in the aiming direction; The aiming display module is used to display the real-time scene image and a preset aiming mark superimposed on the real-time scene image; The target recognition and localization module is used to identify target objects located within the preset aiming mark area in the real-time scene image and generate spatial location information of the target objects; The ultrasonic phased array transducer array is used to generate a parametric array acoustic beam that propagates in a specified direction based on the spatial position information. The audio input and modulation module is used to acquire the audio signal to be transmitted and modulate the audio signal onto the ultrasonic carrier of the parametric array sound beam so that the audio signal can be heard only at the target object.

10. The directional sound transmission device based on vision aiming according to claim 9, characterized in that, The device also includes a target confirmation button and a launch status indicator module. The target confirmation button is used to activate the target recognition and positioning module to lock the target and highlight it on the display screen when the user half-presses the button, and to start transmitting modulated audio beams when the user fully presses the button. The launch status indication module is used to generate a visual prompt on the aiming display module simultaneously when the ultrasonic phased array transducer array launches the parametric array sound beam.