Audio output method, device, electronic device and readable storage medium
By obtaining the location information of the target device and tracking the user's eye gaze information, the audio signal corresponding to the gaze area is automatically output, which solves the problem of cumbersome user operations when the electronic device plays multiple sounds at the same time, and realizes the selective playback of intelligent audio signals.
Patent Information
- Application Number
- CN202210830630.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-15
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-07-15
AI Technical Summary
In a scenario where an electronic device plays multiple sounds simultaneously, it is rather cumbersome for a user to hear a certain sound alone.
By obtaining the location information of the target device, capturing the target image and tracking the user's eye gaze information, the audio signal corresponding to the gaze area is automatically output.
It simplifies user operations and intelligently plays the required audio signals for users, eliminating the need for manual switching.
Smart Images

Figure CN115220683B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to an audio output method, device, electronic device and readable storage medium. Background Art
[0002] Currently, electronic devices can play multiple sounds simultaneously. For example, if two software programs are running simultaneously on an electronic device, one playing a TV series and the other playing a short video, the sound played by the electronic device is a mixture of the two sounds.
[0003] In actual applications, if the user does not want to hear mixed sounds but wants to hear the sound of a specific software, he or she needs to manually close other software or turn off the sound of other software. In particular, if the user needs to frequently switch between playing sounds, he or she needs to frequently close the software or turn off the sound of the software.
[0004] It can be seen that in the prior art, when an electronic device plays multiple sounds simultaneously, the operation is relatively cumbersome when a user wants to listen to a certain sound alone. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to provide an audio output method that can solve the problem in the prior art that when an electronic device plays multiple sounds simultaneously, the user has to perform cumbersome operations when wanting to hear a certain sound alone.
[0006] In a first aspect, an embodiment of the present application provides an audio output method, the method comprising: an electronic device outputting at least two audio signals; obtaining target position information of at least one target device connected to the electronic device; obtaining a target image corresponding to the target position information based on the target position information; when the target image includes a target user, obtaining gaze information of the target user's eyes on the screen of the electronic device; when the gaze information is associated with a first interface displayed on the screen, outputting a first audio signal corresponding to the first interface to the target device, the at least two audio signals including the first audio signal.
[0007] In a second aspect, an embodiment of the present application provides an audio output device, which includes: a first output module for an electronic device to output at least two audio signals; a first acquisition module for acquiring target position information of at least one target device connected to the electronic device; a second acquisition module for acquiring a target image corresponding to the target position information based on the target position information; a third acquisition module for acquiring gaze information of the target user's eyes on the screen of the electronic device when the target image includes a target user; and a second output module for outputting a first audio signal corresponding to a first interface displayed on the screen to the target device when the gaze information is associated with the first interface, the at least two audio signals including the first audio signal.
[0008] In a third aspect, an embodiment of the present application provides an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
[0009] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0010] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method described in the first aspect.
[0011] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the method described in the first aspect.
[0012] Thus, in an embodiment of the present application, the electronic device simultaneously outputs at least two audio signals, which can obtain target position information of at least one target device connected to the electronic device, and then, based on the target position information of the target device, performs image capture at the scene corresponding to the target position to obtain the captured target image, and determines the target user wearing the target device in the target image. Furthermore, when the target image includes the target user, the camera is used to track the target user's eyes to obtain the user's eye gaze information on the screen of the electronic device, and then, based on the gaze information, obtains a first interface that the user's eyes have been gazing at for a long time, that is, the gaze information is associated with the first interface, and then outputs a first audio signal corresponding to the first interface to the target device. It can be seen that based on an embodiment of the present application, the electronic device simultaneously outputs at least two audio signals, and can selectively output the audio signal played in a certain area of the screen to the device worn by the user based on the user's viewing scenario of the area, thereby avoiding manual operation by the user and intelligently playing a certain audio signal for the user to achieve the purpose of simplifying user operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a flowchart of the audio output method according to an embodiment of the present application;
[0014] Figure 2 is a schematic diagram of the positions of the electronic device and the target device according to an embodiment of the present application;
[0015] Figure 3 This is one of the signal diagrams of the embodiment of the present application;
[0016] Figure 4 This is the second signal diagram of the embodiment of the present application;
[0017] Figure 5 is a display schematic diagram of an electronic device according to an embodiment of the present application;
[0018] Figure 6 is a block diagram of an audio output device according to an embodiment of the present application;
[0019] Figure 7 This is one of the hardware structure diagrams of the electronic device according to the embodiment of the present application;
[0020] Figure 8 This is the second hardware structure diagram of the electronic device according to the embodiment of the present application. DETAILED DESCRIPTION
[0021] The following will be combined with the accompanying drawings of the embodiments of the present application to clearly describe the technical solutions of the embodiments of the present application. Obviously, the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.
[0022] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0023] The audio output method provided in the embodiment of the present application may be executed by the audio output device provided in the embodiment of the present application, or an electronic device integrating the audio output device, wherein the audio output device may be implemented in hardware or software.
[0024] The audio output method provided in the embodiment of the present application is described in detail below through specific embodiments and their application scenarios in conjunction with the accompanying drawings.
[0025] Figure 1 A flowchart of an audio output method according to an embodiment of the present application is shown, taking the method applied to an electronic device as an example, including:
[0026] Step 110: The electronic device outputs at least two audio signals.
[0027] Optionally, the electronic device is a mobile phone device, a tablet device, etc.
[0028] The application scenario of this embodiment is, for example, that multiple audio files in an electronic device are running simultaneously, and different audio signals are played respectively.
[0029] For example, live broadcast software and video software run at the same time, one plays live broadcast audio, and the other plays film and television drama audio.
[0030] Step 120: Acquire target location information of at least one target device connected to the electronic device.
[0031] Optionally, the target device is a headphone device or the like.
[0032] Optionally, the electronic device and the target device are connected via a Bluetooth data connection, and correspondingly, Bluetooth data can be transmitted between the electronic device and the target device.
[0033] Furthermore, based on the transmission of Bluetooth data between the two parties, the target device can play the audio signal in the electronic device.
[0034] Optionally, the target location information is relative location information of the target device relative to the electronic device.
[0035] For example, the target location information includes an angle relative to the electronic device. Furthermore, the target location information also includes a distance relative to the electronic device.
[0036] Optionally, when the electronic device is connected to multiple target devices, the target location information of each target device is obtained respectively.
[0037] Step 130: Acquire a target image corresponding to the target position information according to the target position information.
[0038] In this step, based on the acquired target location information, the camera in the electronic device may capture images toward the range indicated by the target location information, that is, capture images of target devices within the range.
[0039] In this step, the captured image is the target image.
[0040] Optionally, when the electronic device is connected to multiple target devices, each user wearing each target device watches the screen of the electronic device at the same time, so that each user can be within the acquisition range of the camera of the electronic device, thereby ensuring that the camera can capture images based on the location information of each target.
[0041] Optionally, the camera is a front-facing camera.
[0042] Optionally, when the electronic device is connected to multiple target devices, target images corresponding to respective target position information are acquired respectively.
[0043] Step 140: When the target image includes the target user, obtain gaze information of the target user's eyes on the screen of the electronic device.
[0044] Based on the target image captured by the electronic device, a target user included in the target image may be determined, where the target user is a user wearing the target device.
[0045] The acquisition range of the camera at least covers the face area of the target user wearing the target device, which is convenient for determining the target user in the target image and for tracking the eyes in the face area of the camera.
[0046] Correspondingly, in this step, when the target image includes the target user, the eyes of the target user can be tracked to achieve the purpose of tracking the target user's line of sight, thereby obtaining the gaze information of the target user's eyes on the screen of the electronic device.
[0047] Optionally, when the electronic device is connected to multiple target devices, gaze information of the target user corresponding to each target device on the screen of the electronic device is obtained respectively.
[0048] Step 150: When the gaze information is associated with a first interface displayed on the screen, output a first audio signal corresponding to the first interface to the target device, wherein the at least two audio signals include the first audio signal.
[0049] Optionally, the first audio signal of the at least two audio signals is played and emitted by the first interface on the screen. Therefore, if it is detected that the target user's eyes are watching the first interface based on tracking the target user's eyes, the first audio signal corresponding to the first interface can be output separately to the target device corresponding to the target user.
[0050] Optionally, when the electronic device is connected to multiple target devices, the audio signal corresponding to the interface viewed by the target user is output to each target device respectively.
[0051] Optionally, in this step, the gaze information is associated with the first interface displayed on the screen, and it is assumed that the user's eyes are more focused on the first interface.
[0052] Thus, in an embodiment of the present application, the electronic device simultaneously outputs at least two audio signals, which can obtain target position information of at least one target device connected to the electronic device, and then, based on the target position information of the target device, performs image capture at the scene corresponding to the target position to obtain the captured target image, and determines the target user wearing the target device in the target image. Furthermore, when the target image includes the target user, the camera is used to track the target user's eyes to obtain the user's eye gaze information on the screen of the electronic device, and then, based on the gaze information, obtains a first interface that the user's eyes have been gazing at for a long time, that is, the gaze information is associated with the first interface, and then outputs a first audio signal corresponding to the first interface to the target device. It can be seen that based on an embodiment of the present application, the electronic device simultaneously outputs at least two audio signals, and can selectively output the audio signal played in a certain area of the screen to the device worn by the user based on the user's viewing scenario of the area, thereby avoiding manual operation by the user and intelligently playing a certain audio signal for the user to achieve the purpose of simplifying user operation.
[0053] In the process of the audio output method of another embodiment of the present application, step 120 includes:
[0054] Sub-step A1: Send a first ultra-wideband (UWB) signal.
[0055] Sub-step A2: When the target device receives the first UWB signal, receive a second UWB signal sent by the target device based on the first UWB signal.
[0056] Optionally, the electronic device includes a first UWB module, the target device includes a second UWB module, and UWB signals can be transmitted between the first UWB module and the second UWB module.
[0057] Therefore, electronic devices can send and receive UWB signals; similarly, target devices can send and receive UWB signals.
[0058] In this embodiment, the electronic device sends a first UWB signal, the target device receives the first UWB signal, and feeds back a second UWB signal to the electronic device based on the first UWB signal, so that the electronic device receives the fed-back second UWB signal.
[0059] Sub-step A3: Acquire target location information of the target device according to the second UWB signal.
[0060] See also Figure 2 , the electronic device 201 can obtain the angle and distance of the target device 202 or the target device 203 relative to the electronic device 201 as the target position information.
[0061] In this embodiment, based on the mutual transmission of UWB signals between the electronic device and the target device, the electronic device can determine the angle and distance of the target device relative to the electronic device through the UWB signal, so that it can accurately capture images of the target device, so as to facilitate subsequent tracking of the eyes of the target user wearing the target device.
[0062] In the process of the audio output method of another embodiment of the present application, the target image is a video image; accordingly, before step 140, the method further includes:
[0063] Step B1: When the video image includes a face sub-image, a vital sign signal corresponding to the face sub-image is obtained according to the video image.
[0064] Step B2: When the vital sign signal indicates a human vital sign signal, determine a target user corresponding to the face sub-image.
[0065] In this embodiment, the images captured by the camera are continuous, and all the captured frame images can form a video. Furthermore, after a facial sub-image is recognized in the video, the vital sign signal corresponding to the facial sub-image can be obtained based on the subsequent video content.
[0066] In some scenarios, the facial sub-images appearing in the captured video may not be real human faces, but rather non-real human faces in images. Therefore, in the captured video, the corresponding life signature signals can be extracted from the facial sub-images, and the life signature signals can be used for biometric recognition to determine whether the extracted life signature signals are indicative of human life signature signals.
[0067] For example, determine whether the vital signs signal has an obvious pulse pattern. Figure 3 The signal shown has obvious changing rules and is a real human life characteristic signal; Figure 4 The signal change pattern shown is not obvious and is not a real human life characteristic signal, such as environmental noise.
[0068] Wherein, a facial sub-image may be identified in the collected image according to facial features, so as to determine the user corresponding to the facial sub-image as the target user.
[0069] In this embodiment, by capturing an image, a face can be identified at the target device based on the target location information. The face can then be further determined to be a real person's face or a face in an image. Consequently, if a real person's face is identified at the target device, the face is identified as the target user's face, and the target user's eye gaze is tracked. This embodiment can thus relatively accurately identify the target user in the target image, ensuring the normal execution of subsequent steps.
[0070] In the process of the audio output method of another embodiment of the present application, step 140 includes:
[0071] Sub-step C1: Obtain an eye tracking heat map formed by the target user's eyes gazing at the screen of the electronic device within the target duration.
[0072] Among them, the eye tracking heat map is used to reflect the gaze location on the screen and the gaze duration corresponding to the gaze location.
[0073] Optionally, the target duration is set to facilitate tracking of gazes within a period of time. Since the gazes at a certain moment are accidental, this can ensure that the obtained tracking results can reflect the real viewing scenario.
[0074] In this embodiment, the camera performs eye tracking on the face of the target user to obtain an eye tracking heat map formed by the target user's eyes based on their gaze on the screen.
[0075] The eye tracking heat map shows all the points on the screen where the target user's gaze is focused, defined as gaze points, and the duration of time spent at each gaze point, defined as gaze duration. For example, the longer the user remains at a gaze point, the darker the color of that gaze point will be. Based on the eye tracking heat map, we can analyze which area of the screen the user's eyes were looking at during the target duration.
[0076] Optionally, the screen is divided into multiple areas, each area is used to display an interface; further, each interface correspondingly sends an audio signal.
[0077] For example, see Figure 5 , the screen simultaneously displays the playback interface 501 of the first video and the playback interface 502 of the second video, and the gaze position point 503 in the figure is a gaze position point in the eye tracking heat map.
[0078] Optionally, the camera determines the angle and distance of the eyes relative to the electronic device based on the target position information and the eye features in the face sub-image, thereby accurately tracking the eyes of the target user.
[0079] In this embodiment, the camera tracks the target user's eyes and can identify the area on the screen where the target user is looking over a certain period of time, thereby controlling the target device worn by the target user to play the audio signal. This embodiment intelligently selects the audio signal to be played for the user to meet their viewing needs.
[0080] In the process of the audio output method of another embodiment of the present application, before step 150, the method further includes at least one of the following:
[0081] Step D1: when the number of gaze position points located within the first interface obtained according to the gaze information is greater than a first threshold, determining that the gaze information is associated with the first interface.
[0082] In this step, one way to determine the association relationship of the gaze information is to find an area on the screen where gaze positions are densely distributed, and then determine the interface corresponding to the area as the interface associated with the gaze information.
[0083] In a certain area, if the number of gaze location points is greater than a first threshold, it is considered that the gaze location points are densely distributed in the area.
[0084] For example, the first threshold may be customized by the system; or, in another example, the first threshold may be manually defined by the user.
[0085] Step D2: when the gaze time of the first interface obtained according to the gaze information is greater than a second threshold, determining that the gaze information is associated with the first interface.
[0086] In this step, another way to determine the association relationship of the gaze information is to find the distribution area with darker color of the gaze position point on the screen, and then determine the interface corresponding to the area as the interface associated with the gaze information.
[0087] In this case, within a certain area, if the time that the eyes view the area is greater than a second threshold, the color of the gaze position point may be darker to a certain extent.
[0088] For example, the second threshold may be customized by the system; or, in another example, the second threshold may be manually defined by the user.
[0089] Optionally, different gaze durations correspond to different background colors.
[0090] Optionally, the above two steps can be used in combination, that is, find a certain area where there are more gaze position points and the color of the gaze position points is darker, and use the interface corresponding to the area as the first interface.
[0091] In this embodiment, when multiple interfaces are displayed on the screen and all are playing audio signals, the associated interface can be found based on the target user's sub-gaze information on each interface and used as the first interface. The first audio signal corresponding to the first interface can then be output to the target device. This shows that based on this embodiment, the appropriate audio signal can be accurately selected for the user based on the user's eye gaze on the screen, providing intelligent services to the user.
[0092] In summary, the purpose of this application is to provide a method for automatically switching audio for an audio device based on eye tracking and UWB signals. The method is mainly based on the use of eye tracking, and the combination of a facial image vital sign algorithm and a UWB angle detection method to determine which video the user's eyes are watching, so as to automatically switch the audio on the audio device to the video they are watching. Both the audio device and the electronic device are equipped with UWB modules, so that the relative angle and relative distance of the audio device are detected based on the UWB signal. Furthermore, based on the image acquisition of the audio device, the vital sign signal of the facial image is extracted for biometric detection. If the biometric recognition is successful, the relative angle and relative distance of the eyes in the facial image are combined with the UWB relative angle and distance to track the line of sight. The tracking results are processed according to the image processing algorithm to obtain the user's eye tracking heat map. Finally, based on the eye tracking heat map, it is determined which video the user's eyes are watching.
[0093] In this application, not only can the audio be automatically switched for the user, but in the scenario where multiple people share the video resources of a single device, the appropriate audio can be matched to the device worn by each user, so that multiple people do not interfere with each other.
[0094] The audio output method provided in the embodiment of the present application can be executed by an audio output device. In the embodiment of the present application, the audio output device provided in the embodiment of the present application is described by taking the audio output method executed by the audio output device as an example.
[0095] Figure 6 A block diagram of an audio output device according to another embodiment of the present application is shown, the device comprising:
[0096] The first output module 10 is used for the electronic device to output at least two audio signals;
[0097] A first acquisition module 20 is configured to acquire target location information of at least one target device connected to the electronic device;
[0098] The second acquisition module 30 is used to acquire a target image corresponding to the target position information according to the target position information;
[0099] a third acquisition module 40, configured to acquire gaze information of the target user's eyes on the screen of the electronic device when the target image includes the target user;
[0100] The second output module 50 is configured to output a first audio signal corresponding to the first interface to the target device when the gaze information is associated with the first interface displayed on the screen, wherein the at least two audio signals include the first audio signal.
[0101] Thus, in an embodiment of the present application, the electronic device simultaneously outputs at least two audio signals, which can obtain target position information of at least one target device connected to the electronic device, and then, based on the target position information of the target device, performs image capture at the scene corresponding to the target position to obtain the captured target image, and determines the target user wearing the target device in the target image. Furthermore, when the target image includes the target user, the camera is used to track the target user's eyes to obtain the user's eye gaze information on the screen of the electronic device, and then, based on the gaze information, obtains a first interface that the user's eyes have been gazing at for a long time, that is, the gaze information is associated with the first interface, and then outputs a first audio signal corresponding to the first interface to the target device. It can be seen that based on an embodiment of the present application, the electronic device simultaneously outputs at least two audio signals, and can selectively output the audio signal played in a certain area of the screen to the device worn by the user based on the user's viewing scenario of the area, thereby avoiding manual operation by the user and intelligently playing a certain audio signal for the user to achieve the purpose of simplifying user operation.
[0102] Optionally, the first acquisition module 20 includes:
[0103] A sending unit, configured to send a first UWB signal;
[0104] a receiving unit, configured to receive, when the target device receives the first UWB signal, a second UWB signal sent by the target device based on the first UWB signal;
[0105] The first acquiring unit is configured to acquire target location information of the target device according to the second UWB signal.
[0106] Optionally, the target image is a video image; the device further includes:
[0107] a fourth acquisition module, configured to acquire, based on the video image, a vital sign signal corresponding to the face sub-image when the video image includes the face sub-image;
[0108] The first determination module is configured to determine a target user corresponding to the face sub-image when the vital sign signal indicates a human vital sign signal.
[0109] Optionally, the third acquisition module 40 includes:
[0110] a second acquiring unit, configured to acquire an eye tracking heat map formed by the target user's gaze on the screen of the electronic device within a target duration;
[0111] Among them, the eye tracking heat map is used to reflect the gaze location on the screen and the gaze duration corresponding to the gaze location.
[0112] Optionally, the device further comprises at least one of the following:
[0113] a second determining module, configured to determine that the gaze information is associated with the first interface if the number of gaze position points located within the first interface obtained according to the gaze information is greater than a first threshold;
[0114] The third determining module is configured to determine that the gaze information is associated with the first interface when a gaze duration of the first interface obtained according to the gaze information is greater than a second threshold.
[0115] The audio output device in the embodiment of the present application can be an electronic device or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices other than a terminal. For example, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc. It can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine or a self-service machine, etc., and the embodiment of the present application does not specifically limit it.
[0116] The audio output device of the embodiment of the present application may be a device having an action system. The action system may be an Android action system, an iOS action system, or other possible action systems, which are not specifically limited in the embodiment of the present application.
[0117] The audio output device provided in the embodiment of the present application can implement each process implemented in the above method embodiment. To avoid repetition, they will not be described here.
[0118] Alternatively, as Figure 7 As shown, an embodiment of the present application also provides an electronic device 100, including a processor 101, a memory 102, and a program or instruction stored in the memory 102 and executable on the processor 101. When the program or instruction is executed by the processor 101, each step of any of the above-mentioned audio output method embodiments is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0119] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.
[0120] Figure 8 A schematic diagram of the hardware structure of an electronic device implementing an embodiment of the present application.
[0121] The electronic device 1000 includes but is not limited to components such as a radio frequency unit 1001 , a network module 1002 , an audio output unit 1003 , an input unit 1004 , a sensor 1005 , a display unit 1006 , a user input unit 1007 , an interface unit 1008 , a memory 1009 , and a processor 1010 .
[0122] Those skilled in the art will understand that the electronic device 1000 may also include a power source (such as a battery) to power each component, and the power source may be logically connected to the processor 1010 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system. Figure 8 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be repeated here.
[0123] Among them, the processor 1010 is used for the electronic device to output at least two audio signals; obtain target position information of at least one target device connected to the electronic device; obtain a target image corresponding to the target position information based on the target position information; when the target image includes a target user, obtain the gaze information of the target user's eyes on the screen of the electronic device; when the gaze information is associated with a first interface displayed on the screen, output a first audio signal corresponding to the first interface to the target device, and the at least two audio signals include the first audio signal.
[0124] Thus, in an embodiment of the present application, the electronic device simultaneously outputs at least two audio signals, which can obtain target position information of at least one target device connected to the electronic device, and then, based on the target position information of the target device, performs image capture at the scene corresponding to the target position to obtain the captured target image, and determines the target user wearing the target device in the target image. Furthermore, when the target image includes the target user, the camera is used to track the target user's eyes to obtain the user's eye gaze information on the screen of the electronic device, and then, based on the gaze information, obtains a first interface that the user's eyes have been gazing at for a long time, that is, the gaze information is associated with the first interface, and then outputs a first audio signal corresponding to the first interface to the target device. It can be seen that based on an embodiment of the present application, the electronic device simultaneously outputs at least two audio signals, and can selectively output the audio signal played in a certain area of the screen to the device worn by the user based on the user's viewing scenario of the area, thereby avoiding manual operation by the user and intelligently playing a certain audio signal for the user to achieve the purpose of simplifying user operation.
[0125] Optionally, the processor 1010 is further configured to send a first UWB signal; when the target device receives the first UWB signal, receive a second UWB signal sent by the target device based on the first UWB signal; and obtain target location information of the target device according to the second UWB signal.
[0126] Optionally, the target image is a video image; the processor 1010 is further used to obtain a vital sign signal corresponding to the facial sub-image based on the video image when the video image includes a facial sub-image; and determine the target user corresponding to the facial sub-image when the vital sign signal indicates a human vital sign signal.
[0127] Optionally, the processor 1010 is further used to obtain an eye tracking heat map formed by the target user's eyes based on their gaze on the screen of the electronic device within a target duration; wherein the eye tracking heat map is used to reflect the gaze position point of the eyes on the screen, and the gaze duration corresponding to the gaze position point.
[0128] Optionally, the processor 1010 is further used to determine that the gaze information is associated with the first interface when the number of gaze position points located within the first interface obtained according to the gaze information is greater than a first threshold; and to determine that the gaze information is associated with the first interface when the gaze duration of gazing at the first interface obtained according to the gaze information is greater than a second threshold.
[0129] In summary, the purpose of this application is to provide a method for automatically switching audio for an audio device based on eye tracking and UWB signals. The method is mainly based on the use of eye tracking, and the combination of a facial image vital sign algorithm and a UWB angle detection method to determine which video the user's eyes are watching, so as to automatically switch the audio on the audio device to the video they are watching. Both the audio device and the electronic device are equipped with UWB modules, so that the relative angle and relative distance of the audio device are detected based on the UWB signal. Furthermore, based on the image acquisition of the audio device, the vital sign signal of the facial image is extracted for biometric detection. If the biometric recognition is successful, the relative angle and relative distance of the eyes in the facial image are combined with the UWB relative angle and distance to track the line of sight. The tracking results are processed according to the image processing algorithm to obtain the user's eye tracking heat map. Finally, based on the eye tracking heat map, it is determined which video the user's eyes are watching.
[0130] In this application, not only can the audio be automatically switched for the user, but in the scenario where multiple people share the video resources of a single device, the appropriate audio can be matched to the device worn by each user, so that multiple people do not interfere with each other.
[0131] It should be understood that in an embodiment of the present application, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042, and the graphics processor 10041 processes the image data of a static picture or video image obtained by an image capture device (such as a camera) in a video image capture mode or an image capture mode. The display unit 1006 may include a display panel 10061, and the display panel 10061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 1007 includes a touch panel 10071 and at least one of the other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include two parts: a touch detection device and a touch controller. Other input devices 10072 may include but are not limited to a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and an action stick, which will not be repeated here. The memory 1009 can be used to store software programs and various data, including but not limited to applications and action systems. The processor 1010 may integrate an application processor and a modem processor, wherein the application processor mainly processes the action system, user pages and applications, etc., and the modem processor mainly processes wireless communications. It is understandable that the modem processor may not be integrated into the processor 1010.
[0132] The memory 1009 can be used to store software programs and various data. The memory 1009 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, applications or instructions required for at least one function (such as a sound playback function, an image playback function, etc.). In addition, the memory 1009 may include a volatile memory or a non-volatile memory, or the memory 1009 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM), and a direct memory bus random access memory (DRRAM). The memory 1009 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0133] Processor 1010 may include one or more processing units. Optionally, processor 1010 integrates an application processor and a modem processor. The application processor primarily handles operations related to the operating system, user interface, and application programs, while the modem processor primarily processes wireless communication signals, such as a baseband processor. It is understood that the modem processor may not be integrated into processor 1010.
[0134] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the above-mentioned audio output method embodiment are implemented and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0135] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0136] An embodiment of the present application further provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-mentioned audio output method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0137] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0138] An embodiment of the present application provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the above-mentioned audio output method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0139] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, it should be noted that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in the opposite order according to the functions involved. For example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0140] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0141] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. An audio output method, characterized in that: The method comprises: The electronic device outputs at least two audio signals; Acquire target position information of at least one target device connected to the electronic device; wherein the target device is a headphone device, and the target position information is relative position information of the target device relative to the electronic device; According to the target position information, acquiring a target image corresponding to the target position information; In a case where the target image includes a target user, obtaining gaze information of the target user's eyes on the screen of the electronic device; the target user is a user wearing the target device; In a case where the gaze information is associated with a first interface displayed on the screen, a first audio signal corresponding to the first interface is output to the target device, and the at least two audio signals include the first audio signal.
2. The method according to claim 1, characterized in that The acquiring target location information of at least one target device connected to the electronic device includes: sending a first UWB signal; In a case where the target device receives the first UWB signal, receiving a second UWB signal sent by the target device based on the first UWB signal; The target location information of the target device is acquired according to the second UWB signal.
3. The method according to claim 1, characterized in that The target image is a video image; and when the target image includes a target user, before obtaining the gaze information of the target user's eyes on the screen of the electronic device, the method further includes: In a case where the video image includes a face sub-image, obtaining a vital sign signal corresponding to the face sub-image according to the video image; In a case where the vital sign signal indicates a human vital sign signal, the target user corresponding to the face sub-image is determined.
4. The method according to claim 1, wherein The obtaining of the gaze information of the target user's eyes on the screen of the electronic device includes: Obtaining an eye tracking heat map formed by the target user's eyes gazing at the screen of the electronic device within a target duration; The eye tracking heat map is used to reflect the gaze position of the eye on the screen and the gaze duration corresponding to the gaze position.
5. The method according to claim 1, wherein In the case where the gaze information is associated with the first interface displayed on the screen, before outputting the first audio signal corresponding to the first interface to the target device, the method further includes at least one of the following: If the number of gaze position points located within the first interface obtained according to the gaze information is greater than a first threshold, determining that the gaze information is associated with the first interface; If the gaze time of the first interface obtained according to the gaze information is greater than a second threshold, it is determined that the gaze information is associated with the first interface.
6. An audio output device, characterized in that: The device comprises: A first output module, the electronic device outputs at least two audio signals; A first acquisition module is configured to acquire target position information of at least one target device connected to the electronic device; wherein the target device is a headphone device, and the target position information is relative position information of the target device relative to the electronic device; A second acquisition module is used to acquire a target image corresponding to the target position information according to the target position information; a third acquisition module, configured to acquire, when the target image includes a target user, gaze information of the target user on the screen of the electronic device; the target user is a user wearing the target device; The second output module is configured to output a first audio signal corresponding to a first interface displayed on the screen to the target device when the gaze information is associated with the first interface, the at least two audio signals including the first audio signal.
7. The device according to claim 6, characterized in that The first acquisition module includes: A sending unit, configured to send a first UWB signal; a receiving unit, configured to receive, when the target device receives the first UWB signal, a second UWB signal sent by the target device based on the first UWB signal; The first acquiring unit is configured to acquire target location information of the target device according to the second UWB signal.
8. The device according to claim 6, characterized in that The target image is a video image; the device further includes: a fourth acquisition module, configured to acquire, based on the video image, a vital sign signal corresponding to the face sub-image when the video image includes the face sub-image; The first determination module is configured to determine the target user corresponding to the face sub-image when the vital sign signal indicates a human vital sign signal.
9. The device according to claim 6, characterized in that The third acquisition module includes: a second acquiring unit, configured to acquire an eye tracking heat map formed by the target user's gaze on the screen of the electronic device within a target duration; The eye tracking heat map is used to reflect the gaze position of the eye on the screen and the gaze duration corresponding to the gaze position.
10. The device according to claim 6, characterized in that The device further comprises at least one of the following: a second determining module, configured to determine that the gaze information is associated with the first interface if a number of gaze position points located within the first interface obtained according to the gaze information is greater than a first threshold; A third determining module is configured to determine that the gaze information is associated with the first interface if a gaze duration of the first interface obtained according to the gaze information is greater than a second threshold.
11. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the audio output method according to any one of claims 1 to 5 are implemented.
12. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the audio output method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Information processing method and electronic device
CN105872371A
Audio and video playback system, playback method and playback device
WO2021238550A1